Forex & Brokerage

How to A/B Test an AI Voice Calling Campaign

Teodor AvadaniTeodor Avadani, Founder·
·11 min read·Last updated:
Cover Image for How to A/B Test an AI Voice Calling Campaign

Checking the dashboard on day two and switching everyone to the arm that's ahead is how brokerages roll out coin flips. AI call campaign A/B testing fixes that with a plan written before the first dial: one variable, two arms drawn from the same dormant segment, a sample size from your own baseline, and a decision rule everyone signed. The maths behind the discipline is blunt: check a 5% significance test after every result and it produces false positives 26.1% of the time.

This guide walks through the seven decisions in order, using the numbers a retention lead at a forex or CFD broker actually has: dormant book size, connect rate, positive outcome rate, deposits. Topcalls runs both arms at $0.35 per minute all-inclusive, and the test plan template linked below is the one-page sheet the method hangs on.

Key Takeaways

  • AI call campaign A/B testing changes exactly one variable between arms; the AI disclosure, recording notice, opt-out line and suppression rules stay fixed in both.
  • Checking results after every call and stopping at the first significant read turns a 5% false-positive rate into 26.1%, per Evan Miller's analysis of the peeking problem.
  • Assign traders to arms by account ID before the first dial; running A in the morning and B in the afternoon makes call window the hidden variable.
  • Under 47 CFR 64.1200, US telephone solicitations stay between 8 a.m. and 9 p.m. at the called party's location, so a call-window test has the same legal ceiling in both arms.
  • On Topcalls, 2,000 test calls averaging 1.5 billed minutes cost about $1,050 at $0.35 per minute all-inclusive, with no per-seat or setup fees.

1. What is AI call campaign A/B testing for a brokerage?

AI call campaign A/B testing runs two versions of an outbound calling campaign against the same segment of traders at the same time, with one difference between them, and reads a single pre-agreed metric at the end. The control is the campaign as it runs today. The variant changes one thing: the opening line, the offer, the call window, the voice, or the retry cadence. Nothing else moves.

Brokers get this wrong because a calling campaign has more moving parts than a landing page. The dormant list changes daily as KYC statuses and balances update, the compliance desk edits a line in one script and forgets the other, and the account managers who take handoffs are on shift for one arm only. Each is a second variable. A test with two variables has no result.

The payoff is that a reactivation campaign stops being a matter of opinion. A test earns its minutes when the winning arm can roll out to the rest of the book; if the honest answer to 'what will we do with a win' is 'nothing', pick a different variable.

2. Which variable should a forex broker test first?

Test the variable that a number you already have is pointing at. Connect rate fine at 38% but conversations ending inside 20 seconds points at the opening line. Long conversations that end in 'not interested' point at the offer. Voicemail on most attempts points at the call window or the days between attempts. Start with the cheapest read: opening-line tests settle on positive outcome rate within days.

VariableExample variantReads onTime to readCompliance check
Opening lineNames the trader's last instrument instead of a generic greetingPositive outcome rateDaysBoth scripts approved as financial promotions
Reason to returnPlatform update walkthrough instead of a check-inReactivation or deposit rate2 to 4 weeksNo incentive that breaches ESMA's CFD restrictions
Call window18:00 to 20:00 local instead of 10:00 to 12:00Connect rateAbout a weekInside legal calling hours in every jurisdiction
Voice or languageNative-language voice for the trader's countryConversation length, outcome rate1 to 2 weeksDisclosure line translated and approved
Retry cadenceSecond attempt after 3 days instead of the next dayConnect rate per trader, opt-out rate2 to 3 weeksAttempt caps and DNC rules unchanged
Variables worth testing in a brokerage AI call campaign

Offer tests carry a trap in the EU and UK. ESMA's 2018 CFD measures include a restriction on incentives to trade, and the same notice cites that 74-89% of retail accounts lose money. So the reason to return in a variant can be a platform walkthrough, a KYC refresh, or an account manager callback. Never a trading bonus. The reactivation offer guide covers what's left once bonuses are off the table.

Brokerage campaign lead comparing two AI call campaign arms on side-by-side dashboards

Each variant script is a financial promotion in its own right. FCA COBS 4.2.1 R requires that a communication or financial promotion is 'fair, clear and not misleading', so the variant gets the same sign-off as the control before launch. The dormant trader call scripts post has openings that already passed that review.

Call-window tests have a legal ceiling in both arms. Under 47 CFR 64.1200(c)(1), no telephone solicitation to a US residential subscriber goes out 'before the hour of 8 a.m. or after 9 p.m. (local time at the called party's location)'. Inside that window there's plenty to test, and the best time to call dormant traders post gives starting hypotheses by market.

3. How do you split dormant traders into control and variant?

Assign each trader to an arm before the first dial, by account ID, with a rule unrelated to list order or time of day. Odd and even account IDs work. So does a random number written into a CRM field. Both arms draw from the same segment: same dormancy window, funded status, jurisdiction and last platform, whether that's MT4, MT5 or cTrader. The split is the part most teams get wrong.

Alternating rows leak the list's sort order, which is usually last-login date or balance. Running A this week and B next week leaks a market event or a platform outage. Splitting A into the morning and B into the afternoon turns call window into the variable, and you didn't mean it to. Every one of these feels random from the inside and isn't.

On Topcalls the clean shape is two campaigns with the test ID in each name, loaded from two lists split by account ID in the CRM export, so outcome tags and exports join later on the test ID. Same voice, same disclosure, same retry rules on both, except the one row that's the test.

Scrub suppression and DNC on the whole segment before splitting, not on each arm after. A number on the UK TPS register or a national do-not-call list comes out of both lists, and so does anyone who opted out on a previous campaign. Under 47 CFR 64.1200(d)(3) a US do-not-call request must be honored within 10 business days, and a test is not a reason to miss that.

4. How many calls does an AI call A/B test need?

Size the test from your own baseline, not from a vendor benchmark. Pull the primary metric from the last 30 days of the control, decide the smallest lift you'd act on, pick a confidence level, and put those three numbers into a standard two-proportion sample size calculator. The output is calls per arm. Divide by attempts per trader to get traders per arm, then check the segment fills both arms after suppression.

The minimum lift is where the honesty lives. Moving positive outcome rate from 12% to 15% is a lift most retention leads would roll out, and detecting it needs far fewer calls than a move from 12% to 12.5%. Pick the lift you'd act on, not the one that makes the test short. A test that ends with 9 positive outcomes in one arm and 11 in the other tells you nothing, whatever the percentages look like.

Then price it: total calls, times average billed minutes, times the per-minute rate. Topcalls charges $0.35 per minute all-inclusive, with no per-seat or setup fees, so 2,000 calls averaging 1.5 billed minutes come to about $1,050. Set that against what one reactivated trader is worth; the dormant trader revenue calculator does that arithmetic from your own deposit and volume numbers.

If the segment can't fill both arms, widen the dormancy window, raise the minimum lift, or run the test longer. Don't shrink the arms and hope. And if 'today' isn't a stable campaign yet, run the control alone for a week first; a scoped reactivation pilot is the right place to get that baseline.

The test plan template is one page: the hypothesis line, a control-versus-variant table with a 'held constant' column, sample size fields, a metric and guardrail table, stopping rules, a decision table and a running test log.

5. Which metric decides an AI calling test?

One primary metric decides the test, and it's the fastest metric that still reflects money. Opening-line and call-window tests read on positive outcome rate or connect rate within days. Offer tests need reactivation or deposit rate and a longer attribution window. Guardrails, such as opt-out rate, complaint count and handoff queue wait, can pause an arm or veto a rollout, but they never pick the winner.

Broker retention, compliance and ops leads reviewing AI call A/B test results before rollout

Fix the outcome tags the AI agent may use before launch and keep them identical in both arms: 'will fund', 'wants callback', 'KYC help needed', 'not interested', 'do not call again'. If the variant's tag list differs, the primary metric isn't comparable. Count the attribution window from each trader's last attempt, not the campaign start date, and keep it identical for both arms.

Where the numbers live matters as much as what they are. One sheet, one owner, both arms side by side, exports from the calling platform and the trading platform joined on account ID and test ID. Topcalls' real-time analytics show connect rate, outcome tags and billed minutes per campaign; the deposit side comes from MT4, MT5 or the back office.

A variant that wins on outcomes and doubles opt-outs doesn't roll out. Set every guardrail threshold from the control's history, so the pause rule is a number, not a feeling. The conversion rate guide lists the metrics calling teams track and which deserve a guardrail.

6. When do you stop an AI call campaign A/B test?

Stop when both arms reach the sample size from the plan and the end date has passed, and not before. Cover at least one full week per arm so weekday and weekend windows appear in both. The only earlier stops are a guardrail breach, which pauses the breaching arm the same day, and an operational failure, such as a disclosure missing from one arm or a leaked assignment, which invalidates the arm outright.

The reason for the rule is arithmetic, not patience. Evan Miller worked it through in 'How Not To Run an A/B Test': with a fixed sample size and a 5% significance level you expect 5% false positives, but if you check after every observation and stop at the first significant read, the false positive rate climbs to 26.1%. His instruction: 'Decide on a sample size in advance and wait until the experiment is over before you start believing the "chance of beating original" figures.'

So keep the primary metric split out of the daily view. A daily guardrail read is fine. The A-versus-B outcome rate goes in a sheet the team opens once, at the end. And write a freeze list: scripts, offer, call windows, attempt rules and retry cadence, suppression files, voice, handoff routing. If something has to change mid-test, both arms change on the same day and the change is logged.

7. What does a broker do with the test result?

Read the result against a decision table agreed before launch. A variant that beats the control by at least the minimum lift with no guardrail breach rolls out. One that loses or ties means the control stays and the finding is logged. A win on the primary metric with a guardrail breach means no rollout, plus a note on why. The table makes the review meeting about rollout mechanics, not interpretation.

Rollout on Topcalls is a campaign edit, so name the person with access, list the segments the winner applies to, and set the day the old control switches off. Then fill one line in the test log: test ID, variable and segment, result versus control, decision, rolled-out date. Twelve lines of that log beat any vendor's benchmark, because they're your book, your segments and your traders.

Every test produces the next hypothesis, and the loser's call recordings usually say why it lost. An opening line naming the trader's last instrument that still lost may have lost because the instrument data was six months stale: a data hygiene test, not a script test. Write the next candidate while the result is fresh. For a second pair of eyes on which variable to run first, book a 30-minute call and bring your last 30 days of outcome tags.

8. When doesn't A/B testing an AI call campaign fit?

A/B testing doesn't fit when the segment is too small to fill two arms, when the campaign itself isn't stable yet, or when the change is a compliance fix rather than a hypothesis. A dormant segment of 400 accounts after suppression can't detect a three-point lift in any reasonable window. Fix those for everyone, immediately, and don't test them.

It also strains at a broker with three markets and a different script per language, where every arm splits again by country and none reach sample. Run one market first, or scope the work as a pilot. And skip the test when the answer wouldn't change anything: if compliance has already ruled the variant offer out, or ops can't roll a winner to the rest of the book, spend the minutes on the control.

One variable, two arms, a sample size from your own baseline, and a decision table signed before launch. That's the whole method. The template puts it on one page.

Frequently Asked Questions

Get AI calling tips in your inbox

No spam. One email per week with actionable sales automation tips.

Share this article

XLinkedIn

Summarize with AI

Ready to automate your calls?

Book a 30-min call or calculate your ROI.

Related Articles