What does it mean to split test a call script?
Split testing a call script means running two versions of the same call on the same list at the same time. You keep the one that books more.
Half your dials hear opener A. Half hear opener B. Everything else stays identical, so the only thing that can explain a gap in results is the words.
Marketers have done this with ads and subject lines for twenty years. Almost nobody does it with phone calls, because with human callers it is nearly impossible.
Two salespeople never deliver a script the same way twice. One had a bad morning, the other had a coffee. The test is contaminated before dial ten.
An AI voice agent removes that problem completely. Variant A is delivered identically on every single call. So is variant B. For the first time, a small NZ or AU business can run a properly controlled test. That kind of rigour used to need a 50 seat call centre and a statistician.
Why do opinions about scripts cost you bookings?
Because the confident opinion usually wins the meeting, and the meeting is not where bookings happen.
Every sales team has a script argument running. Should the agent mention price early or late? Should the opener lead with the caller's suburb or the offer? Should it ask a question in the first five seconds or state the reason for the call?
These arguments settle by seniority in most businesses. The owner picks the version they would respond to. Then a thousand dials run on a guess.
Here is the uncomfortable part. On campaigns we run, wording changes have shifted booking rates by a third. Not the offer. Not the list. The words.
If the loser ships because the wrong person won the argument, you pay for it on every dial. The call costs the same either way, roughly 80 cents a minute billed by the second.
How does a split test actually run on a live campaign?
The mechanics are simple, and they matter less than the discipline around them.
Then you keep the winner and go again with a new challenger. The best script you will ever run is not written in a strategy session. It is found, one honest test at a time.
What should you test first?
Test the parts of the call where most hangups happen, in this order.
The first ten seconds. More calls die here than everywhere else combined. Test leading with the reason for the call against leading with a question. Test naming the caller's area against not naming it.
On NZ and AU mobiles, only 47 to 65 percent of dials connect at all. Those first seconds decide whether the rest of your script ever gets heard.
The disclosure. Every Waboom AI agent tells the caller it is an AI at the start of the call. That is not negotiable.
How the sentence lands is testable. Plain and quick against warm and explained. You keep the honesty, you tune the delivery.
The ask. Booking a specific slot against offering a choice of two. Asking for the appraisal against asking for ten minutes.
Small changes here move the last step of the funnel. That is the step you invoice against, as we covered in cost per booked outcome.
The objection lines. When the caller says not right now, the next sentence decides whether that contact stays warm.
Test a soft callback offer against a simple thanks and exit. Then check the follow up data. A brushed off caller is not the same as an interested one, and your agent should treat them differently.
What sample size do you need before the result means anything?
More than feels natural, fewer than you fear.
A gap of two bookings after 40 dials is noise. Judge the early funnel first: connects, survival past the opener, and conversation length all read clearly at a few hundred calls per variant. Booking verdicts need more patience.
The textbook maths says a small three point lift in bookings takes over two thousand conversations per variant to prove. So hunt landslides in your results, not decimals. A variant that doubles how long callers stay past the opener is a signal. A variant that books 8 percent more over 60 conversations is a coin flip.
The good news is that volume arrives fast when an agent is dialling. A campaign doing 100 dials a day puts real numbers on the board inside two weeks.
And because the agent delivers each variant identically, every conversation is a clean data point. No caller mood. No Friday afternoon slump.
If your volumes are small, run the test longer rather than calling it early. The maths does not care that you are impatient.
One more piece of honesty the industry skips. No voice vendor anywhere has published a properly quantified split test of script variants. The category ships the split without the statistics. We would rather hand you the real maths.
What does a split test cost to run?
Nothing extra. That is the part most operators miss.
You are already paying for the dials, at about 80 cents a minute billed by the second. A split test does not add calls. It adds structure to calls you were making anyway. The only real cost is the discipline to write a second variant and the patience to let the test finish.
Compare that with what the winner is worth. Take a campaign doing 1,000 dials a month at a booking rate of 1.8 percent, so 18 booked outcomes. A winning script that lifts bookings by a third takes that to 24.
Same list, same spend, six more outcomes every month. It compounds for as long as the campaign runs. On the campaign economics we published, that is the cheapest growth you will buy this year.
When should you stop testing?
Never fully, but slow down once the gains do.
Early tests find big wins because early scripts carry big flaws. After four or five honest rounds, the gaps narrow. That is the signal to shift attention from wording to the other levers: list quality, time windows, and follow up rules. A script that has survived five challengers is a strong script. At that point the next test matters less than what happens to the leads after the call.
Keep one habit permanently. Any time you change the offer, the market, or the season, the old winner is a challenger again. A script that won in a spring property market has not proven anything about winter.
Frequently Asked Questions
Can I split test inbound calls too?
Yes, though the variable is usually the greeting and the routing question rather than an opener. Randomise by call rather than by list, let each variant take a few hundred calls, and compare booked outcomes and transfer rates the same way.
Does the AI disclose itself in both variants?
Always. Every Waboom AI agent states it is an AI at the start of the call, in both variants of every test. Disclosure wording can be tuned. Disclosure itself is fixed. Do Not Call suppression runs across every campaign and is never part of a test.
What data do I see per variant?
Dials, connects, conversation length, outcome tags and bookings for each variant, side by side in the portal. Call transcripts stay attached to each record, so when a variant wins you can read exactly why. Transcripts and structured call data sit on our Sydney servers, while live audio is processed offshore. We are transparent about that split.
How is this different from just rewriting my script when results dip?
A rewrite replaces your script with a new guess. A split test makes the new guess earn the job against the current holder, on the same list at the same time.
The difference sounds small and is everything. One is opinion with better timing. The other is evidence.
Leonardo Garcia-Curtis
Founder & CEO at Waboom AI. Building voice AI agents that convert.
Ready to Build Your AI Voice Agent?
Let's discuss how Waboom AI can help automate your customer conversations.
Book a Free Demo


