Forex & Brokerage

How to Run Quality Assurance for AI Voice Calls

Teodor AvadaniTeodor Avadani, Founder·
·11 min read·Last updated:
Cover Image for How to Run Quality Assurance for AI Voice Calls

Most brokerages listen to zero AI calls after launch. AI voice call quality assurance is the habit that fixes that: a fixed sample of recordings each week, scored against the same seven checks, so you know whether the agent is losing dormant traders in the opening line, the objection handling, or the CRM write-back.

This guide covers which calls to pull, how to score each part of the conversation, what the FCA, the EU AI Act and the FCC expect from the recording and the disclosure, and how to turn a week of findings into one script fix.

Key Takeaways

  • A weekly QA pass on 10 to 20 AI call recordings per campaign, scored on seven checks, tells a brokerage which single fix to make before the next batch dials.
  • FCA rule SYSC 10A.1.14R requires relevant telephone recordings to be kept for five years, and up to seven if the FCA asks.
  • Article 50(1) of the EU AI Act, applying from 2 August 2026, requires people to be told they're interacting with an AI system unless that's already obvious.
  • The FCC ruled on February 8, 2024 that AI-generated voices are "artificial" under the TCPA, so US disclosure and consent checks belong in every QA review.
  • Topcalls includes recording, transcription and analytics in its $0.35 per minute price, so the raw material for QA costs nothing extra.

1. What does AI voice call quality assurance actually check?

AI voice call quality assurance is a weekly review of recorded AI calls against a fixed list of checks: the opening, whether the agent understood the trader, how it handled interruptions, what it did with objections, the handoff to a human, the AI and recording disclosures, and whether the logged outcome matches the call. Each check gets a pass, a fail or not applicable, with a timestamp as evidence.

That's a different job from campaign monitoring, which watches connect rate, average duration and dispositions per hour in real time. QA listens to the calls behind those numbers and asks why. A campaign can post a healthy connect rate and still be burning dormant traders with an opening that sounds like a robocall.

The seven checks map to the seven places an AI agent goes wrong on a trader call. A KYC reminder that sounds like a sales pitch is an opening problem. Answering "how do I deposit?" when the trader asked "do I still have money in the account?" is a comprehension problem. Logging "not interested" when the trader said "call me after payday" is a disposition problem, and that one quietly deletes your follow-up pipeline.

The checklist behind this article puts those seven checks on one printable page with a sampling section in front. Reviewers of Topcalls AI voice agents use the same split, script, voice, routing or CRM mapping, because each has a different owner.

2. Which AI calls should a brokerage review each week?

Pull a fixed number of recordings per campaign every week, split across every disposition your CRM uses, plus every call handed to a human, every call tied to a trader complaint, every call under 20 seconds and every call over your expected maximum length. Ten to twenty sampled calls per campaign is enough to spot a pattern; handoff and complaint calls are reviewed in full.

Reviewing only the funded or booked calls is the most common QA mistake. The "not interested" and dropped calls tell you why the script fails, and they're where the compliance findings hide.

Brokerage QA reviewer scoring an AI voice call recording against a printed checklist

Cover every language and voice the campaign runs. Topcalls places calls in 32 languages, and a script that reads naturally in English can come out stiff in Arabic. One call per language per week is the floor. Write the script version on each review so you can tell later whether a fix helped.

Open the recording and the transcript side by side: the transcript shows what the speech recognition heard, the recording shows what the trader said. Account numbers, deposit amounts and platform names like MT4, MT5 and cTrader are where recognition slips most, and every gap between the two is a finding. The sampling rules are the first section of the QA checklist, so the same kind of calls get picked whoever is reviewing.

3. How do you score the opening and comprehension?

Score the opening on five things in the first ten seconds: the agent waited for the trader to speak, named the brokerage and the reason for the call in the first sentence, used the trader's name correctly, referenced something true about the account, and replied without a pause long enough for a second "hello?". Comprehension is scored on whether the agent answered the question that was actually asked.

Time the first reply. Topcalls targets sub-500ms response latency, and a reviewer can hear the difference between that and a two-second gap without a stopwatch. If several sampled calls have a long pause at the top, that's a platform finding, not a script one, and it belongs with the vendor. The sibling guide on AI voice latency for outbound calls covers how to test it before you sign.

Then check what the agent referenced. "You opened an MT5 account in March and the last trade was in June" tells the trader this is their broker. "I'm calling with an exciting opportunity" could be anyone. Count how many sampled calls have the trader asking "who is this?" inside 30 seconds. More than a couple means the opening isn't doing its job, which is a script fix; the dormant-trader call script guide has openers that pass this test.

Comprehension fails look like the agent answering a nearby question. Mark every place the transcript differs from the audio, check the agent followed a mid-call language switch (traders in Dubai or Kuala Lumpur do this often), and confirm it asked for clarification rather than guessing wrong on "close my account" versus "call me back". The serious one: any balance, bonus term or spread the agent said that isn't in the campaign data. Escalate that the same day.

4. How should QA treat interruptions, objections and handoffs?

Count interruptions and score the ratio: how many times the trader spoke over the agent, and how many of those the agent stopped for within a word or two. Match every objection to a branch in your objection flow and check the agent stopped after the second no. For handoffs, confirm the agent offered a human the first time the trader asked, and that the human had the context on screen.

Interruption handling is where AI calls either feel natural or feel like a phone tree. The checks are mechanical: did the agent stop mid-sentence rather than finishing it, did it deal with what the trader said before returning to the script, and did it avoid restarting from the top after a cough or an "mm-hm" that wasn't a real interruption. What good sounds like is covered in why interruption handling makes AI calls feel natural.

Objections carry the compliance risk. An agent that responds to "I lost money last time" with a promise of returns, a bonus term it wasn't given, or a spread claim has just created a financial-promotions problem. Check that "remove me from your list" was treated as an opt-out and confirmed, and that a stuck withdrawal or an open support ticket was routed to a human rather than argued with.

CheckWhat you hear on the recordingOwner of the fix
OpeningAgent talks over "hello"; trader asks "who is this?"Script
ComprehensionAnswers a nearby question; misses MT4/MT5 or amountScript or knowledge base
InterruptionsFinishes its sentence, then restarts from the topVoice platform
ObjectionsPromises returns or a bonus term it wasn't givenScript and compliance
HandoffTrader repeats everything to the humanRouting and CRM mapping
Common QA fails on AI calls to traders and who fixes them

Handoff review is short but strict. Traders should be told who they'll speak to and roughly how long the wait is, the human should see the trader's name, account, campaign and what the AI already covered, and if no human was free a callback should have been booked with a specific time and actually placed. When a call should move to a person at all is covered in when an AI call should hand off to a human agent.

5. What do regulators expect from AI call recordings and disclosures?

Three rules shape the disclosure and recording section of a broker's QA review. In the UK, FCA rule SYSC 10A.1.14R requires relevant telephone recordings to be kept for five years, and up to seven if the FCA asks. Article 50(1) of the EU AI Act requires people to be told they're interacting with an AI system. And the FCC treats AI-generated voices as "artificial" under the TCPA, so US consent rules apply.

Brokerage operations team reviewing weekly AI call QA results on a dashboard

The FCA's SYSC 10A rules go further than storage. SYSC 10A.1.16R makes firms responsible for "the quality, accuracy and completeness of the records of all telephone recordings and electronic communications" and for keeping them "readily accessible and available to clients, upon request." For a QA reviewer that means checking the recording exists for every sampled call, covers the whole conversation including the human part, and matches the transcript. A recording that cuts out at the transfer is a finding. Whether your reactivation calls fall inside the recording scope at all is a question for your compliance desk.

The EU AI Act's Article 50(1) requires AI systems that interact directly with people to be built so that "the natural persons concerned are informed that they are interacting with an AI system," unless that's obvious to a reasonably well-informed person. Those transparency obligations apply from 2 August 2026. So the QA check is simple: did the agent identify itself as an AI or automated assistant in the opening, in the approved wording, in the trader's language, and did it answer "is this a real person?" truthfully.

On February 8, 2024 the FCC announced a unanimous Declaratory Ruling recognising that "calls made with AI-generated voices are 'artificial' under the Telephone Consumer Protection Act," which its press release says holds AI voices "to those same standards," including prior express written consent for telemarketing. FCC Chairwoman Jessica Rosenworcel said: "We're putting the fraudsters behind these robocalls on notice." For a brokerage with US traders, the reviewer confirms the identification and callback route in every call and voicemail, and logs any missing or altered disclosure for counsel the same day; the TCPA guide for AI calling has the consent detail. Topcalls builds the disclosure line and recording notice into the campaign on the secure infrastructure side, so the reviewer checks that it fired rather than writing it.

6. How do you turn QA findings into a script fix?

Total the fails per section across the week's sample. The section with the most fails is what you fix next, and each section has one owner: opening, comprehension and objections go to the script, interruptions and latency go to the voice platform, handoff and disposition go to routing and CRM mapping. Change one thing, then re-review the same section on the next batch to see whether it moved.

Disposition accuracy is the check most teams skip and the one that costs the most money. "Callback requested" needs a time, a channel and an owner attached, or the lead dies. "Not interested" where the trader raised a fixable problem should carry that problem as a reason code. Opt-outs and wrong numbers must reach the suppression list. And the call summary must match the transcript, with nothing the trader didn't say.

Put a number on it. If your reactivation campaign produces 60 "callback requested" outcomes a week and QA finds a fifth of them have no time or owner, that's 12 warm traders a week nobody calls back. Run the dormant trader revenue calculator with your average first redeposit to see what that leak costs over a quarter. The calls themselves are cheap: Topcalls charges $0.35 per minute all-inclusive with recording and transcription in the price, so the QA sample costs reviewer time and nothing else.

Score on Monday, ship the one fix by Wednesday, and let the real-time analytics show whether the number moved by Friday. Pair the checklist with the AI call quality metrics post if you want a scorecard with weights. Want a second pair of ears on your first week of recordings? Book a 30-minute call and we'll review a sample with you.

7. When this doesn't fit

Weekly QA on sampled recordings doesn't fit a campaign that isn't recording, a list too small to sample, or a team with nobody who can own the fix. It also doesn't replace a compliance review of the script before launch: QA finds what went wrong on real calls, it can't approve wording in advance.

No recordings, no QA. If consent or jurisdiction means calls aren't recorded, you're limited to transcript and disposition checks, which catch comprehension and CRM problems but not tone, talk-over or latency.

Under a few hundred calls a month. A one-person retention desk working 200 dormant traders by hand can listen to their own calls. Formal sampling adds paperwork without signal until volume grows.

Nobody owns the script. Findings nobody can act on are a weekly complaint log. Name the script owner, the routing owner and the CRM owner before the first review.

Pre-launch approval. Compliance sign-off on the opening line, the disclosure wording and the objection branches happens during customer reactivation setup, before the first call.

Ten calls a week, seven checks, one fix. That's the whole method, for a dormant-trader campaign in Cyprus or a KYC reminder run out of Dubai.

Frequently Asked Questions

Get AI calling tips in your inbox

No spam. One email per week with actionable sales automation tips.

Share this article

XLinkedIn

Summarize with AI

Ready to automate your calls?

Book a 30-min call or calculate your ROI.

Related Articles