Waboom AI
AI Training

AI Training

AI Team TrainingPopular

Hands-on workshops for marketing, sales, operations, and customer service teams.

AI Strategy Workshop

Executive workshops for leadership teams. Identify opportunities. Calculate ROI. Walk out with a roadmap.

Claude Code Workshop

Build apps in hours not months. Ship websites, automations, and tools with AI.

AI Training for Teams

AI Training for Teams

Hands-on workshops for marketing, sales, operations, and customer service teams. Not theory. Real tools. Real tasks. Real outcomes.

2,000+ people trained across NZ

Learn more
AI Automation

AI Automation

AI Agents & AutomationPopular

Your AI workforce: outbound, proposals, knowledge and support agents. Find buyers, write SOWs, answer every call.

AI Retainer Support

Already built with us? Stay on retainer and we keep shipping new agents and features for your business.

Microsoft Copilot Agents

Build custom Copilot agents in Power Automate & Copilot Studio. Automate workflows across your entire Microsoft 365 ecosystem.

Waboom Concierge

Personalised inbound for premium brands. An AI concierge greets every visitor, builds an on-the-spot quote, and books a real conversation.

AI Automation & Integration

AI Automation & Integration

We build faster and more cost effectively than traditional development teams. You tell us the problem. We deliver the solution.

30+ projects live in 24 months

Learn more
AI Voice Agents

AI Voice Agents

AI Voice Agents

24/7 AI-powered phone agents for inbound & outbound calls. Never miss a lead, handle enquiries, book appointments automatically.

AI Receptionist

Pay-as-you-go inbound receptionist. Answers, transfers calls, takes messages inside your VoIP. $1/min with auto top-up.

Voice Agent Pricing

Transparent pricing for AI voice agents. See costs per minute and platform fees.

AI Voice Agent Demo

Talk to Michelle on three voice AI engines side by side. Hear the latency, find the model that fits.

Listen to Our Voices

Preview all 32 AI voice agents across NZ, AU, UK and US. Find the perfect voice for your brand.

Case Studies

Real customer results. Vendor leads, viewings booked, relationships scaled. Every story has the math.

AI Voice Agents

AI Voice Agents

Never miss a lead. AI agents that answer calls 24/7, qualify prospects, and book appointments automatically.

30+ voice agents deployed

Learn more
Case Studies

Case Studies

Melbourne: 5 listings from one 14-year dormant contactPopular

$2.9M of CBD apartments relisted by the same agent who sold them in 2012. AI dialled the dormant number.

Home builder: AU$374.4M in lost sales uncovered

5,200 cold calls into a 70,000-prospect CRM. 234 confirmed lost deals at AU$1.6M each. A very leaky bucket.

Sydney agent: 141 vendor leads in 90 days

9,856 dials, 1,997 conversations, 141 warm-transferred sellers at $32.74 each.

Christchurch developer: 49 viewings in 14 days

931 Meta leads called same-day. 49 viewings booked at $7.12 each.

City Sales Auckland: 100,000+ relationships

How a leading Auckland firm strengthened over 100,000 client relationships with AI.

See all case studies

Browse every Waboom customer case study in one place.

Real numbers from real Waboom customers

Real numbers from real Waboom customers

Vendor leads. Viewings booked. Relationships scaled. Every story has the math.

5,000+ AI-handled conversations

Learn more
Resources

Resources

AI Resources & Guides

Free templates, frameworks, and implementation guides to help you adopt AI effectively in your organisation.

Blog

Expert insights on AI voice agents, automation strategies, and industry best practices from the Waboom team.

Workshop Tutorial Videos

Paid-attendee video library. Cowork 101 and Claude Code 101, on demand with clickable chapter navigation.

AI Resources Hub

AI Resources Hub

Free tools and guides to help you implement AI effectively. From policy templates to ROI calculators.

New resources added monthly

Learn more
Contact

Contact

Contact Us

Get in touch with our team. We'd love to hear about your AI goals.

About Waboom AI

Learn about our mission, team, and why we're passionate about AI adoption in NZ.

Let's Talk AI

Let's Talk AI

Whether you need training, automation, or strategy - we're here to help you adopt AI effectively.

Response within 24 hours

Learn more
09 885 9695 (NZ)+61 485 027 479 (AU)
Back to BlogPerformance

Why Is My Voice Agent's Latency So High on Phone Calls?

Leonardo Garcia-Curtis06/08/2025
TL;DR

Voice AI latency is the six layer stack between your caller's last word and your agent's first: speech to text, LLM processing, knowledge base retrieval, function calls, text to speech, and network round trip. Unoptimised they stack to 1,000 to 3,500ms. A Tauranga commercial cleaning prospect walked out of our demo when the agent paused three seconds. That was a 2,800ms agent 14 months ago. Today our deployments respond in under 800ms end to end. This guide walks every lever we pulled to close the gap, with the exact numbers we hit on each layer.

Why Is My Voice Agent's Latency So High on Phone Calls?

I was on a call with a prospect in Tauranga. Commercial cleaning company, 8 staff. They wanted AI to handle inbound booking enquiries.

Halfway through the demo, the voice agent paused. Two seconds of silence. Then three.

"Is it broken?" she asked.

It wasn't broken. The LLM was thinking. But that 3-second gap killed the demo.

She'd already decided it wasn't ready. We lost the deal.

That was 14 months ago. The agent's response time was 2,800ms. Today, our agents respond in under 800ms. Here's every optimisation that got us there.

The Latency Stack: Where Your Seconds Go

When someone talks to your AI voice agent, here's what happens between their last word and the agent's first:

1. Speech-to-text (STT). Converting audio to text. Typically 20-50ms. Not your problem.

2. LLM processing. The AI thinking about what to say. This is the big one: 500-900ms for most models. GPT-4o sits around 300-600ms.

Smaller models hit 200ms but trade off quality.

3. Knowledge base retrieval. If your agent looks up documents to answer a question, add 50-100ms per lookup.

4. Function calls. Your CRM lookups and calendar checks add 200-2,000ms.

5. Text-to-speech (TTS). Converting the response to audio. 200-500ms. You'd think this would be faster by now.

6. Network round-trip. Data between your servers and the voice platform. 20-80ms.

Add it all up: you're looking at 1,000-3,500ms for a complete turn. And that's before you've even accounted for the "ghost latency", the gaps that no metric shows you.

We built a diagnostic tool that found 2,412ms of hidden latency in one client's setup. You wouldn't see it in any individual metric.

The total? Brutal. Hidden in the gaps between components.

Voice agent latency breakdown

Your caller doesn't care why it's slow.

Why Latency Kills Your Conversions

A 1-second response feels natural. A 2-second response feels like hold music. A 3-second response? Your caller's already checking their phone.

Here's what we've measured across our deployments:

  • Under 1 second: Your conversion holds. Natural flow.
  • 1-2 seconds: 15% engagement drop.
  • Over 2 seconds? 35% drop. Over 3? Your prospect hangs up.

    Your agent's intelligence doesn't matter if nobody stays long enough to experience it.

    The Optimisations That Actually Work

    1. Choose the Right LLM Tier

    Voice platforms offer different performance tiers. The difference for your callers isn't subtle:

  • Standard tier: Variable performance. Good for your testing phase.
  • Fast tier: Dedicated resources, optimised routing. Sub-600ms.
  • Fast tier is what your production agents should use. No debate.

    The cost difference is negligible compared to the conversion impact. If you're running campaigns that call 500 prospects a day, every 100ms of latency reduction translates to real revenue.

    2. Write Prompts Like You're Paying by the Word

    Long prompts = slow responses. Your LLM processes the entire system prompt on every turn.

    Here's a real example. Your before-and-after speaks for itself:

    Before (430 tokens): "Provide a comprehensive explanation of our return policy including exceptions and the return process."

    After (85 tokens): "Explain returns: 30 days, original packaging, receipt needed."

    Same information. 80% fewer tokens. 40% faster. Your callers notice the speed difference.

    3. Enable Response Streaming

    This is the single biggest win you'll get. Without streaming, your agent waits for the complete response before speaking. With streaming, it talks after the first few words.

    Your caller hears the agent respond in 200-300ms instead of 800-1,200ms. Same total time. Dramatically different experience.

    Turn it on. It's a toggle in your agent settings. No reason not to.

    4. Optimise Your Knowledge Base

    If your agent uses RAG-powered knowledge retrieval, the structure of your documents matters:

  • Clear headings and sections, the retrieval engine finds relevant chunks faster
  • One concept per document, don't dump everything into a single file
  • Limit knowledge bases per agent, each additional KB adds retrieval time
  • Keep documents under 50 pages, chunking works better with focused content
  • A well-structured knowledge base adds 50ms. A poorly structured one adds 300ms+. That gap compounds on every single turn of your conversation.

    5. Minimise Function Call Overhead

    Every external API call your agent makes during a conversation adds latency. Some are unavoidable (checking calendar availability). Others are premature.

    Rule of thumb: Don't call your CRM to look up a customer record until you've confirmed their identity. Don't check inventory until they've expressed interest.

    Pre-load what you can. Cache what doesn't change. Defer what isn't urgent. Every function call you remove saves 200-500ms.

    The Settings That Matter

    Your agent configuration has a few levers that directly impact latency:

    Temperature: Lower values (0.2-0.4) produce more predictable, faster responses. Higher values (0.7+) add creativity but also processing time. For production agents, stay below 0.4.

    Max tokens: Cap your responses at 100-150 tokens for voice. Nobody wants a 60-second monologue from an AI agent. Shorter responses mean faster delivery and more natural turn-taking.

    Interruption sensitivity: Set this right. If your agent keeps talking when the caller tries to interrupt, they'll think it's broken. But too sensitive and it'll stop mid-sentence at every background noise.

    Monitoring: Know Before Your Callers Complain

    Don't wait for complaints. Track these metrics from day one:

    P50 end to end latency. Your typical response time. Target: under 1.2 seconds.

    P90 end to end latency. What your worst 10% of calls experience. Target: under 2.5 seconds.

    First token time. How quickly your agent starts speaking. This drives perceived responsiveness more than total duration.

    If your P90 exceeds 3 seconds, you have a problem. Even if your average looks fine. The caller on the wrong end of that P90 doesn't care about your averages.

    Our latency diagnostic tool breaks down exactly where those seconds go, including the "gap" latency that standard dashboards miss.

    Latency monitoring dashboard

    Our dashboard shows the gap no other platform surfaces.

    Industry-Specific Considerations

    Real estate and cold calling. Speed is everything. Your prospect didn't ask for this call. Sub-second responses keep them engaged. Use fast tier, minimal prompts, no knowledge base lookups unless absolutely necessary.

    Healthcare. Precision matters more than speed here. A slightly slower response that gives accurate medical information beats a fast one that hallucinates. Use structured LLM responses and accept the 200ms trade-off.

    Financial services. Pre-load account data before the greeting. When your agent already knows the caller's name and account status, the first turn feels instant. Use fast tier for compliance checks.

    E-commerce. Cache your product catalogue. If your agent has to query your inventory API on every product question, you're adding 300-800ms. Pre-loaded product data keeps responses under a second.

    The Bottom Line

    Perceived latency matters as much as measured performance. Fast feels trustworthy. Slow feels broken.

    Every millisecond you shave off your agent's response time is a conversion you keep. We went from losing a deal in Tauranga at 2,800ms to closing 30% more calls at 800ms.

    Same agent. Same script. Faster delivery.

    Slow agents cost you deals. We fix that.

    Book a Strategy Call | See the Platform

    Frequently Asked Questions

    What's a good target for AI voice agent response time?

    Under 1 second end to end is ideal. Your P50 should sit below 1.2 seconds, and your P90 below 2.5 seconds.

    Anything over 3 seconds and you'll see measurable drops in engagement and conversion. For cold calling, aim for sub-800ms, your prospect didn't ask for this call, so every millisecond of silence works against you.

    Which latency component has the biggest impact?

    LLM processing time dominates. It typically accounts for 500-900ms of your total response time.

    Optimising your prompts (shorter system instructions, fewer tokens) and using the platform's fast tier are the two highest-ROI changes you can make. After that, focus on function call reduction and knowledge base structure.

    Does response streaming actually make a difference?

    Yes, it's the single biggest improvement you'll get for perceived latency. Without streaming, your caller waits for the complete response.

    With streaming, the agent starts speaking after just 200-300ms. Same total processing time.

    Dramatically faster experience. No downside. Turn it on.

    How do I diagnose latency issues I can't see in the dashboard?

    Standard dashboards show component-level metrics (LLM time, TTS time, STT time). But they miss the gaps between components, the "ghost latency" that adds up.

    We built a latency diagnostic tool that surfaces exactly where your hidden seconds are going. It's found 2,400ms+ of invisible delay in agents that looked "fine" on every individual metric.

    LG

    Leonardo Garcia-Curtis

    Founder & CEO at Waboom AI. Building voice AI agents that convert.

    Ready to Build Your AI Voice Agent?

    Let's discuss how Waboom AI can help automate your customer conversations.

    Book a Free Demo

    Related Pages

    AI Voice Agents

    The complete guide to AI voice agents for New Zealand and Australian businesses.

    AI Voice Agents NZ

    Every New Zealand voice-agent service in one place.

    AI Receptionist NZ

    24/7 inbound call answering with native Kiwi accent.

    Related Articles

    Your Call Containment Rate Is the Number Vendors Inflate. Here Is the One Waboom AI Really Holds.

    Your Call Containment Rate Is the Number Vendors Inflate. Here Is the One Waboom AI Really Holds.

    Still Paying $3 Per Inbound Call? Here Is What an Inbound Call Centre AI Voice Agent Costs Instead

    Still Paying $3 Per Inbound Call? Here Is What an Inbound Call Centre AI Voice Agent Costs Instead

    It's 2am and a Customer Is Ringing. Waboom AI Voice Agent Uptime and Failover Keep Your Line Answered.

    It's 2am and a Customer Is Ringing. Waboom AI Voice Agent Uptime and Failover Keep Your Line Answered.

    Waboom AI

    Empowering New Zealand and Australian businesses with AI voice agents and automation that deliver real, measurable value.

    info@waboom.ai+64 9 885 9695 (NZ)+61 485 027 479 (AU)
    Level 8, 139 Quay Street
    Auckland CBD, New Zealand

    Voice Agents

    • AI Voice Agents
    • AI Receptionist NZ
    • AI Receptionist Australia
    • AI Phone Answering
    • AI Virtual Receptionist
    • AI Receptionist Pay As You Go
    • Waboom Concierge
    • Medical Answering Service
    • Answering Service Australia
    • AI Sales Agent
    • Voice Agent Pricing
    • Listen to Voices
    • Real Estate Guide

    By Industry

    • Real Estate
    • Mortgage Brokers
    • Insurance Brokers
    • Property Managers
    • Medical Clinics
    • Dentists
    • Vets
    • Childcare + ECE
    • Car Dealerships
    • Construction + Builders
    • Electricians
    • Plumbers
    • HVAC
    • Accountants
    • Law Firms
    • All industries and regions

    Workshops

    • All Workshops
    • AI Team Training
    • AI Strategy Workshop
    • AI Champion Workshop
    • Claude Team Training
    • Claude Code Workshop
    • Lovable Workshop
    • Free AI Workshop

    Automation

    • AI Automation
    • Microsoft Copilot Agents
    • Integrations

    Company

    • About Us
    • Contact
    • Partners
    • Pipedrive Partner
    • Resources
    • Blog
    • AI Agency NZ
    • AI Agency Australia

    Powered by leading AI technologies

    VAPIOpenAIZapierMakeStripe

    © 2026 Waboom.ai. All rights reserved.

    PrivacyTermsSecurity