Forex & Brokerage

10 AI Call Quality Metrics Brokerage Teams Should Track

Teodor AvadaniTeodor Avadani, Founder·
·11 min read·Last updated:
Cover Image for 10 AI Call Quality Metrics Brokerage Teams Should Track

A dormant-trader campaign can post a 30% connect rate and still be quietly burning the list. AI call quality metrics are how a brokerage catches that: ten scores per sampled call, graded the same way every week, covering the opening, the disclosure, what the agent heard, what it claimed, and what it wrote back to the CRM.

Campaign dashboards count outcomes. They don't tell you the agent misheard "Thursday" as "Tuesday", quoted a bonus that expired last month, or kept pitching after a trader said stop. This guide covers the ten metrics worth tracking on reactivation, deposit and KYC calls, how to score them on a 1 to 5 scale, and how a weekly review turns "the calls sound fine" into a number compliance will sign.

Key Takeaways

  • Ten AI call quality metrics, each scored 1 to 5 and weighted, give a brokerage a 0 to 100 total per call that two reviewers can reproduce.
  • Grading 20 sampled calls every Monday beats grading 100 once a quarter, and the sample has to include every language and outcome the campaign runs.
  • The FCC ruled on 8 February 2024 that AI-generated voices are "artificial" under the TCPA, so US telemarketing calls with an AI voice need prior express written consent.
  • FCA rule SYSC 10A.1.14R requires relevant call recordings to be kept for five years, and up to seven on request, so every score should be keyed to a recording ID.
  • Disposition accuracy is the metric every campaign number rests on: one wrong tag in ten shifts reactivation rate, handoff rate and cost per reactivated trader by 10%.
  • Two scorers whose weighted totals differ by more than 10 points on the same call need a definitions fix before the trend means anything.

1. What Are AI Call Quality Metrics for a Brokerage?

AI call quality metrics are per-call scores that grade how an AI voice agent handled a conversation with a trader, separate from whether the call converted. For a forex or CFD brokerage that means ten checks: opening and disclosure, right-party verification, comprehension, latency, factual accuracy, compliance guardrails, objection handling, handoff, disposition accuracy and next-step clarity. Each one is scored 1 to 5 from the recording or the transcript.

Campaign metrics and quality metrics answer different questions. Connect rate, reactivation rate and cost per reactivated trader tell you what happened across 5,000 dials. Quality metrics tell you why call 3,412 went wrong, and whether the same fault is sitting in the other 4,999. Reactivation campaign metrics measure the campaign; the ten metrics below measure each call.

The distinction matters more on a finance call than on a pizza order. A misheard deposit amount, a leverage figure quoted to the wrong entity, or a pitch that continues after "don't call me again" is a complaint, and in some jurisdictions a breach. That's why a customer reactivation campaign at a brokerage needs a quality layer the pizza place can skip. If you're setting up the review process itself, the AI voice call quality assurance guide covers sampling and cadence; this article covers what to score.

2. Which Five Metrics Show the Conversation Worked?

The first five AI call quality metrics grade the mechanics of the conversation: whether the agent opened with the brokerage name and disclosure, verified it was speaking with the account holder, understood what the trader said, kept the turn-taking natural, and got its facts right. Factual accuracy carries the heaviest weight in this group, because a wrong number on a finance call is a complaint waiting to happen.

  • Opening and disclosure: A 5 means brokerage name, reason for calling and the required disclosure all land before the trader's first reply. A 1 is the trader asking "who is this?" twenty seconds in. Article 50(1) of the EU AI Act, applying from 2 August 2026, requires that people interacting with an AI system are told so unless it's obvious from the circumstances. Score the line as delivered, not as written in the prompt.
  • Right-party verification: A name check or "am I speaking with the account holder?" answered yes before any account detail. Telling whoever picked up that the account has been dormant since March is a 1, and on most compliance desks an auto-fail.
  • Comprehension accuracy: Count misheard turns in the transcript. Zero is a 5. One misunderstanding the agent caught and repaired ("sorry, did you say Thursday?") is a 3. A callback booked for the wrong day is a 1.
  • Latency and turn-taking: ITU-T Recommendation G.114 says most applications are unaffected when one-way delay stays under 150 ms and sets 400 ms as the ceiling for network planning. An AI agent adds listening and thinking time on top of the network, and Topcalls keeps that full response cycle under 500 ms. On the recording, score the gaps: a trader saying "hello? hello?" into silence is a 2, and an agent that talks over the trader is a 2 as well.
  • Factual accuracy: Every claim on the call checked against the knowledge base and the trader's entity terms. "The 20% bonus is still active" when it expired last month is a 1, whatever else went right. Multi-entity brokers should check the leverage cap quoted against the regulator that entity sits under.
Brokerage QA reviewer scoring an AI call transcript against a ten-metric scorecard

The AI Call Quality Scorecard packs all ten metrics into one sheet: a definition, a way to measure it, a weight and a 1 to 5 scale for each, plus seven auto-fail rules and a weekly review layout your QA lead can fill in on a Monday.

3. Which Five Metrics Prove the Call Did Its Job Safely?

Metrics six to ten grade outcome and safety: compliance guardrails, objection handling, handoff and escalation, disposition accuracy, and next-step clarity. Together they answer whether the call moved the trader forward without saying anything the compliance desk would later have to explain. Disposition accuracy sits in this group because every campaign number, from reactivation rate to cost per reactivated trader, is built on that one CRM tag.

  • Compliance guardrails: No forbidden phrasing anywhere, risk wording where your policy requires it, and a stop request honoured in one sentence. The FCC's Declaratory Ruling of 8 February 2024 recognises calls made with AI-generated voices as "artificial" under the TCPA, which puts US telemarketing calls with an AI voice under the prior express written consent rule. FCC Chairwoman Jessica Rosenworcel said in the announcement: "We're putting the fraudsters behind these robocalls on notice." Score the consent basis as part of this metric, not as an afterthought.
  • Objection handling: Dormant traders push back with "I lost money", "the spreads were bad" or "I moved brokers". A 5 acknowledges the objection, answers it from the approved list and offers a next step. Reading the same rebuttal twice, or arguing, is a 2.
  • Handoff and escalation: Score both directions. A trader who asked for a person, raised a complaint or mentioned a withdrawal problem and stayed with the agent is a missed handoff. A transfer on a routine "what are your hours?" is a wasted one. The human handoff workflow for AI calls lists the triggers worth wiring.
  • Disposition accuracy: Compare the CRM tag to the transcript. "Callback requested, Thursday 2pm" matching a real request is a 5. "Not interested" on a call where the trader asked for the KYC link is a 1, and it also just cost you a reactivation. Call disposition automation is how the tag gets written without a human in the loop, which is exactly why it needs checking.
  • Next-step clarity: The agent states the next step, the trader confirms it, and the CRM shows a matching task or a link sent. A next step mentioned but never confirmed is a 3. No next step on a positive call is a 1.

4. How Do You Score AI Call Quality Consistently?

Score each metric 1 to 5 from what actually happened on the recording, not from what the prompt says should happen. Multiply each score by its weight, add them up and divide by 5 for a 0 to 100 total. Mark any metric the call never reached as N/A and drop its weight from the denominator. Then write one sentence of evidence with a transcript timestamp for every score below 4.

#MetricWhat a 5 looks likeWhere to check
1Opening and disclosureName, reason and disclosure before the first replyFirst 15 seconds of the recording
2Right-party verificationHolder confirmed before any account detailTranscript, before first account mention
3Comprehension accuracyZero misheard turnsTranscript, agent replies vs trader turns
4Latency and turn-takingNo dead air, no talking over the traderRecording, gaps between turns
5Factual accuracyEvery claim matches the knowledge base and entity termsTranscript vs knowledge base
6Compliance guardrailsNo forbidden phrases, stop requests honouredTranscript vs forbidden-phrase list
7Objection handlingAcknowledged, answered, next step offeredTranscript, objection turns
8Handoff and escalationRight calls transferred, routine ones keptTranscript plus transfer log
9Disposition accuracyCRM tag matches what the trader saidCRM record vs transcript
10Next-step clarityStated, confirmed, and visible in the CRMEnd of transcript plus CRM task
The ten AI call quality metrics at a glance

Two people grade the first ten calls independently each week. If their weighted totals differ by more than 10 points on any call, they listen together and rewrite the metric definition that caused the split. Keep one reference recording per score level for the three heaviest metrics, so a new scorer can calibrate in an hour instead of a month.

Auto-fails sit outside the scale. Guaranteed-return language, personal trading advice, account details shared before verification, continuing after a stop request, a missing recording or AI disclosure, deposit pressure on a trader who mentioned financial difficulty, or a bonus quoted to the wrong entity: any one of these zeroes the call and goes to compliance the same day.

"Accuracy felt off" is not evidence. "At 0:42 the agent said the bonus was still active; it expired last month" is, and it's fixable by Tuesday.

5. How Do Quality Scores Connect to Campaign Revenue?

Brokerage operations and compliance team reviewing weekly AI call quality scores by language

Quality scores explain campaign numbers. Say reactivation rate drops from 6% to 4% between two weeks: on a dashboard that's a mystery, on the scorecard it's a diagnosis, because comprehension fell in one language or objection handling slipped after a prompt edit. Topcalls records and transcribes every call inside its $0.35 per minute price, so the evidence behind each score costs reviewer time and nothing else.

Disposition accuracy is the bridge. Cost per reactivated trader divides campaign spend by the count of "reactivated" tags, so a 10% error rate in tagging moves that number by 10% before anyone touches a script. Run the dormant trader revenue calculator with your own deposit and volume figures, then ask what a tagging error of that size does to the result. It's usually larger than the prompt tweak everyone wanted to argue about.

Real-time analytics on Topcalls shows outcomes per campaign, language and segment as calls land. The scorecard adds the per-call grade a dashboard can't infer. Put the two side by side in the weekly review and the "why" column stops being a guess. For cross-industry reference points, the AI cold calling metrics and benchmarks post covers connect and conversion ranges outside brokerage.

6. What Should the Weekly Quality Review Cover?

A weekly review takes about 60 minutes and covers six things: 20 calls sampled at random and split across outcomes and languages, the lowest scoring call and its cause, the average gap between the two scorers, the metric that moved most and the change that explains it, fixes shipped since last week with the metric each one targets, and any segment or language below the flag threshold. The QA lead signs; compliance countersigns any week with an auto-fail.

Force the sample split. A quarter each of positive outcomes, negative outcomes, human handoffs and voicemails, with every language the campaign dials represented. A campaign averaging 78 overall with one language at 58 has a problem the average hides. Live campaign monitoring catches drops and cost pacing during the day; the scorecard catches what the agent actually said.

Keep the signed sheets with the recording IDs. FCA rule SYSC 10A.1.14R requires firms to keep relevant telephone recordings for five years, and up to seven where the FCA asks, so a score that can't be traced to a recording ID is an opinion, not a record. The same logic applies under MiFID II for EU entities and under your own retention policy everywhere else.

Tie every prompt or knowledge base change to a metric on the sheet. A change without a target metric is a guess, and a metric that moved with no change behind it usually means the sample drifted. Topcalls sets up a first campaign in about 15 minutes, so a fix found on Monday can be dialing on Tuesday; the AI voice agents page covers how a campaign's instructions and knowledge base are set.

7. When This Doesn't Fit

A ten-metric scorecard is too much for a brokerage running under 200 AI calls a week, for a desk where humans still make every outbound call, or for a campaign with no recordings or transcripts to grade from. It also fails without an owner: if nobody on the compliance desk will sign the forbidden-phrase list and the auto-fail rules, the scores never become a record anyone can rely on.

  • Low volume: Under 200 calls a week, listen to every handoff and every complaint in full and skip the weighted total. Sampling adds nothing when you can hear everything.
  • Human-only desks: Grade human agents on the same ten metrics once a quarter for comparison. Weekly scoring of 20 human calls is a coaching program, not a QA program, and needs a different sheet.
  • No transcript access: If your vendor charges extra for recordings or transcripts, fix that before building a scorecard. Topcalls includes both in the $0.35 per minute price. A score without evidence is a mood.
  • Regulator-prescribed formats: Some regulators and internal audit teams prescribe their own call review form. Map the ten metrics onto that form rather than running two reviews in parallel.

Start with the setup section this week: pick the sample size, name two scorers, set the flag and pause thresholds, and get the auto-fail list signed. Grade 20 calls next Monday. If you'd rather see how Topcalls surfaces these metrics on a live broker campaign, book a 30-minute call and bring last week's lowest scoring recording.

Frequently Asked Questions

Get AI calling tips in your inbox

No spam. One email per week with actionable sales automation tips.

Share this article

XLinkedIn

Summarize with AI

Ready to automate your calls?

Book a 30-min call or calculate your ROI.

Related Articles