Home Hub How It Works Features Use Cases How-To Guides Help Docs Pricing Login
AI

Voice AI Hallucinations: How to Catch False Answers Before Customers Hear Them

Radomir Basta • September 26, 2026 • 19 min read
Voice AI Hallucinations Mitigation with Multi AI Platform by Suprmind.

In November 2022, Jake Moffatt used Air Canada’s website chatbot to ask about bereavement fares after a death in the family. The chatbot said the reduced fare could be claimed within 90 days after travel. Moffatt booked at full price and applied. Air Canada refused, because its real policy did not allow retroactive claims.

The detail that matters most: the chatbot pointed Moffatt to Air Canada’s own bereavement travel page, and that page said the opposite. The correct answer was one click away. The bot still got it wrong.

When the case reached British Columbia’s Civil Resolution Tribunal in 2024, Air Canada argued, in effect, that the chatbot was responsible for its own words. The tribunal called that “a remarkable submission”, held the airline liable for negligent misrepresentation and ordered it to pay C$812.02 in damages, interest and fees.

That was a text chatbot, not a voice agent. On a phone call the same error lands harder. There is no link to click, no transcript to scroll back through, no second tab where the customer might spot the contradiction. They hear the answer once, in a warm and confident voice, and act on it.

More of these calls are coming. In a Gartner survey of 321 customer service leaders, 91% said executive leadership was pushing them to implement AI in 2026.

This guide covers what actually reduces false answers on customer calls. The short version: a voice agent is not one model that needs stricter instructions. It is a chain of systems, and every link in the chain can produce a fluent, confident, wrong sentence.

A voice agent can be wrong in seven different places

The most useful recent evidence comes from τ-Voice, a benchmark published in March 2026 as a voice extension of the τ²-bench agent benchmark. It put voice agents through 278 realistic customer service tasks across retail, airline and telecom: returns, flight changes, plan changes, each with real policies to follow and real backend actions to take.

A text-based reasoning agent completed 85% of the tasks. Voice agents built on realtime APIs from OpenAI, Google and xAI completed 31-51% with clean audio, and 26-38% once background noise and diverse accents were added.

Two findings matter more than the headline numbers.

First, reviewers traced 79-90% of failures to agent behavior, not to bad audio. The agents made reasoning and policy mistakes even when transcription was accurate.

Second, the failures were specific and familiar to anyone who has run a contact center. Authentication was the biggest bottleneck, with agents mistranscribing names and emails even when callers spelled them letter by letter. And in one simulated call, an agent told the customer “I’ve updated your shipping address” without ever calling the tool that updates it.

Read the numbers with the right caveat. The benchmark tested general-purpose realtime models inside a research harness, not production deployments with custom retrieval, validation and handoff built around them, and the authors note that cascaded speech-to-text, LLM and text-to-speech pipelines were not included. The distance between 85% and 31% is the work production engineering exists to close. It is an argument for building the system around the model with care, not an argument against voice agents.

Here is where that system can break:

Where it breaksWhat happenedExampleMain control
HearingSpeech recognition got a word, name or number wrongOrder B3172 becomes B3712Readback and confirmation
RetrievalThe system pulled a wrong, stale or irrelevant documentLast year’s refund policy wins the searchKnowledge base hygiene
GenerationThe model distorted good evidence or filled a gap“Your plan includes unlimited seats”Grounding plus claim checks
Tool callRight intent, wrong function or wrong argumentsCancels the other booking on the accountValidation at the tool
StateThe backend failed but the agent reported success“Your refund is on its way” after an API errorConfirm only from success responses
AuthorityA fluent sentence the agent had no right to sayAn invented discount or goodwill creditPolicy rules outside the prompt
VerificationNothing checked a high-risk claim before it was spokenA wrong balance, read out with confidenceRisk-tiered verification

Only one of these seven is the classic hallucination people picture. The other six are system failures that sound exactly like hallucinations to the caller. The caller does not care which component was at fault. Neither did the tribunal in Moffatt.

Static facts belong in retrieval. Live facts belong in tools.

Retrieval-augmented generation (RAG) is the standard foundation for customer-facing agents, and for good reason. Instead of answering from what the model absorbed in training, the agent searches your approved knowledge base mid-call and answers from what it finds. Done well, it removes a large share of invented answers about policies, products and procedures.

It does not turn the model into a lookup table. Retrieval can return the wrong passage, and the model can misread the right one. A 2025 ACL paper named the first problem hallucination on hallucination: flawed retrieval hands the model a flawed premise, and the answer compounds the error. The Air Canada chatbot illustrates the second. The correct policy was published and even linked, and the answer still contradicted it.

Grounded summarization benchmarks tell the same story at scale: even when the source document is handed straight to the model, the best models on those leaderboards still record nonzero hallucination rates, which our AI hallucination rates and benchmarks page tracks across providers.

In customer service, most retrieval failures start in the knowledge base, not the search algorithm:

  • Two versions of the same policy are both indexed, and the outdated one matches the caller’s wording better.
  • A policy that varies by region, plan or date is stored as one flat document.
  • The exception sits in paragraph six of a PDF that the chunker split in half.
  • Pages nobody owns never get retired.

This is why retrieval design matters as much as retrieval itself. How documents are chunked, versioned and filtered decides what the model gets to see. The fix is partly editorial. Give every knowledge article an owner, a version and a review date. Retire old versions instead of publishing new ones beside them. Store conditional policies as structured rules (region, plan, date range) rather than prose the model has to interpret on the fly. The same Gartner survey found 58% of service leaders plan to upskill agents into knowledge management specialists, which is a fair signal of where the work sits.

The bigger fix is to stop asking the knowledge base questions it cannot answer. A policy document can explain how refunds work. It cannot tell a caller whether their refund was issued. Facts that are customer-specific and change over time need a live query to the system that owns them.

The caller asksSource of truthWhy
“What time do you close on Sunday?”Knowledge baseStatic, the same for everyone
“Where is my order?”Order management APIChanges by the hour
“How much do I owe?”Billing systemCustomer-specific and legally sensitive
“Is Tuesday at 2pm free?”Calendar or booking systemChanges by the minute
“Did you cancel it?”The success response from the cancellation callThe only proof the action happened
“Can you make an exception?”Policy rules engine, or a humanA question of authority, not information

One rule covers most of this table. The agent should never state a customer-specific fact it did not read from an authoritative system during this call. Not from the prompt, not from last month’s call summary, not from what sounds plausible.

Running all of this inside a live conversation is the hard part. Speech recognition, turn-taking, retrieval, function calls and speech output all have to finish inside a response window callers experience as natural. That is why many teams build on an AI voice agent platform that already handles knowledge grounding, mid-call function calling and human handover as one stack, and spend their own engineering time on the part only the business can define: which claims come from which source, and what the agent is allowed to promise.

Never let a mishearing become a database action

A text agent receives exactly what the user typed. A voice agent receives a transcription, which is a best guess about what the caller said. For most words the guess is good enough. For identifiers it is dangerous, because a single wrong character still produces a valid-looking order number, and valid-looking order numbers open the wrong customer’s account.

OpenAI’s current guidance for realtime voice agents is direct about this. It tells developers to treat order IDs, tracking numbers, account numbers, confirmation codes, phone numbers and email addresses as high-precision values, confirm the final value before any account lookup or write action, and ask a short clarifying question when audio is unclear rather than act on a guess.

Here is the difference on a real call:

Caller: It’s B three one seven two.

Weak agent: Thanks, I found order B3712. It shipped yesterday.

Strong agent: I heard B, 3, 1, 7, 2. Is that right? … Thanks. I see two items going to Denver on that order. Is that the one?

The weak agent heard one digit wrong and answered about someone else’s order with full confidence. The strong agent gave the caller two chances to catch the error, and the second check makes a wrong ID fail out loud instead of silently.

Practical techniques that hold up on phone lines:

  • Read identifiers back in short chunks the caller can check.
  • Use a spelling alphabet for letters that collapse on narrowband audio: B and D, M and N, F and S.
  • Validate format server-side before the lookup. If order numbers are one letter and four digits, a six-character string triggers a clarification, not a query.
  • Offer keypad entry for long numbers. For a 16-digit number, DTMF tones are far more reliable than speech recognition.
  • Confirm against a second detail the caller knows, such as the delivery city or the last item ordered.
  • Cap retries. After two failed attempts, switch paths (text a secure link, or transfer to a person) instead of looping. τ-Voice recorded agents going unresponsive after repeated authentication failures.

Guard the actions, not just the words

A wrong sentence is a trust problem. A wrong action is an operations problem: the cancelled booking, the refund to the wrong card, the address change on someone else’s account. Actions need controls the model cannot talk its way past.

Confirm only from a success response. The τ-Voice agent that “updated” an address without calling the tool shows that the agent’s belief about what it did is not evidence. Build confirmation sentences from the API response, not from the model’s recollection. The amount, the card ending, the reference number all come from the system. If there is no success response, the agent says the action did not go through.

Validate at the tool, not in the prompt. Every function should check its own arguments, the caller’s authentication state and the account’s eligibility before it executes. OpenAI’s agent guardrails documentation puts the principle in one line: place validation next to the tool that creates the side effect.

Confirm the consequence before anything irreversible. Read back what will happen, to what, and what it costs: “To confirm, you want to cancel the Friday 7pm booking for four people. There is no cancellation fee. Shall I go ahead?”

Route high-impact actions to approval. Refunds above a threshold, account closures, anything that touches payment details. The threshold is a business decision based on your fraud and error tolerance, not a number from a blog post. Current agent frameworks support pausing a run until a person or a policy approves the tool call.

Authority failures deserve their own list. Write down what the agent may never offer on its own: discounts, compensation, fee waivers, policy exceptions, delivery guarantees, medical or legal guidance. Enforce the list in code wherever you can. A prompt instruction is a request, and a caller who insists “your colleague promised me last week” is running a live social-engineering test on it.

Verify the claims that can hurt you, and let the rest stream

Voice has a constraint text does not: silence. A study of ten languages in PNAS found the most common gap between one speaker finishing and the next one starting falls between zero and 200 milliseconds. Callers notice delays that chat users never would. A full verification pass on every sentence would make an agent accurate and unbearable.

So spend verification where being wrong is expensive. The cost of reasoning should rise with the cost of being wrong.

Risk tierExamplesTreatment
LowGreetings, process guidance, “I can help with that”Stream immediately
MediumPolicy facts from the knowledge baseGrounded in retrieved text, sampled in QA review
HighPrices, balances, dates, eligibility, commitments, anything that triggers an actionChecked against the source before it is spoken

For the high tier, the check is usually mechanical and fast. Extract the numbers, dates and names from the draft response and compare them with the tool output they should come from. If the draft says $49 and the billing API says $94, the sentence never reaches the speaker. Where a judgment call is needed, use a separate verifier rather than asking the same model whether it was right. A model grading its own homework shares its own blind spots.

The pause is cheaper than it looks. The same PNAS research found people take longer to respond when the answer is a no or an “I don’t know”. A brief “let me pull that up” before a balance or a policy exception sounds like a careful person, not a slow machine.

Two popular fixes that do less than you think

Lowering the temperature. Temperature controls how much randomness goes into word choice. At zero, a model gives more consistent answers. Consistent is not the same as correct. An agent at temperature zero that misread your refund policy will misstate it the same way on every call. Low temperature is a sensible setting for support. It is not a factuality control.

Flagging hedge words in transcripts. Scanning calls for “I think”, “maybe” and “probably” catches the agent being cautious, which is often exactly when it is right to be. The costly errors arrive with full confidence. Our Multi-Model AI Divergence Index calls this the Confidence Trap: the gap between how certain an AI sounds and how well its answer holds up when another model reviews it. A better transcript signal is a factual claim with no matching source in the call’s retrieval or tool log.

Test the whole call, not just the model

Evaluations on typed prompts miss most of what breaks on the phone. Test the agent the way customers will actually reach it:

  • Real telephony audio. Narrowband phone lines, mobile compression, a speakerphone in a moving car.
  • The accents and languages your callers actually have, including switching languages mid-sentence if your agent supports it.
  • Noise. Kitchens, traffic, a television, a second person talking in the room.
  • Barge-in. Callers who interrupt, change their minds or correct themselves (“no, the other order”).
  • Bad inputs. Wrong order numbers, closed accounts, two customers with the same name.
  • Backend failure. Slow APIs, timeouts, error responses. This is where claimed-success failures hide.
  • Conflicting knowledge. Two articles that disagree. Does the agent notice, pick one, or escalate?
  • Pressure. Callers who claim a promise was made, or try to argue the agent out of its rules.

The τ-Voice design is worth copying: a simulated caller driven by a capable LLM, with configurable accents, noise and turn-taking behavior, working through scripted tasks with known correct outcomes. Turn every real failure you find into a regression test, and run the suite before every prompt, model or knowledge base change. A fix to the refund flow that quietly breaks address changes is easy to ship and hard to notice without it.

Measure unsupported claims, not tone

MetricWhat it catchesTarget
Unsupported claim rateFactual statements with no matching source in the retrieval or tool logDown
Entity error rateWrong IDs, names, amounts or dates used in lookups or spoken backDown
Wrong tool-call rateRight intent, wrong function or argumentsDown
Claimed-without-success rateThe agent confirmed an action that has no success responseZero. Treat every instance as a bug
Escalation precisionShare of handoffs that genuinely needed a personUp
Missed escalation rateCalls that should have reached a person and did notDown
Caller correction rateHow often callers say “no, that’s wrong” or repeat themselvesDown
Severity-weighted failure rateErrors weighted by cost, so a wrong opening time counts less than a wrong balanceDown

Two of these metrics need one piece of plumbing. Log every retrieval result and tool response alongside the transcript, with timestamps. Then any spoken claim can be checked against what the agent actually knew at the moment it spoke. Without that join, call review measures tone, not truth.

Human handoff is part of the design, not the escape hatch

The appetite for human help has not gone away. In a SurveyMonkey study, 79% of Americans said they strongly prefer a human to an AI agent for customer service, and 89% said companies should always offer the option to reach one.

Treat handoff as a planned outcome with explicit triggers:

  • The caller asks for a person. Once is enough.
  • Identification fails twice.
  • A high-risk claim could not be verified against its source.
  • The request needs an exception, compensation or judgment the agent is not allowed to give.
  • The caller has corrected the agent twice in one call.
  • Tools are failing and the task cannot be completed.

Then make the transfer worth something. Pass the summary, the verified identity, what was attempted, what succeeded, what failed and what is still pending. The customer should never have to repeat an order number to the person who picks up. Handled that way, a handoff reads as good service. Handled badly, it reads as the machine giving up.

If you serve callers in the EU: Article 50

Since 2 August 2026, Article 50 of the EU AI Act requires AI systems that interact directly with people, voice assistants included, to make clear the person is dealing with AI, unless that is obvious from context. The information has to arrive no later than the first interaction, which on a phone call means the greeting. The more natural your agent sounds, the harder it is to argue the AI nature is obvious, so put the disclosure in the opening line.

The Digital Omnibus on AI, published in July 2026, postponed several high-risk deadlines but left Article 50 substantively unchanged. Transparency breaches can draw fines of up to 15 million euros or 3% of worldwide annual turnover, whichever is higher. The duty to design disclosure into the system sits with the provider, so if you license a voice agent, check your contract for which party handles which obligation. This is general information, not legal advice.

Pre-launch checklist

  • Every claim type maps to a source of truth: knowledge base, live API, rules engine or human.
  • Knowledge articles have owners, versions and review dates, and old versions are retired rather than left indexed.
  • High-precision entities are read back and confirmed before lookups and writes.
  • Unclear audio triggers a clarifying question, never a guess.
  • Action confirmations are built from success responses.
  • Tools validate their own arguments, authentication and eligibility.
  • Irreversible and high-value actions confirm the consequence or wait for approval.
  • The never-promise list is enforced outside the prompt.
  • High-risk claims are checked against source before they are spoken.
  • Tests run on real telephony audio with noise, accents, barge-in and backend failures.
  • Retrieval and tool logs are joined to transcripts for claim-level review.
  • Handoff triggers are explicit, and every transfer carries full context.
  • The greeting tells callers they are talking to AI.

What building verification for five AIs taught us

Suprmind is not a voice platform. It is a multi-AI chat platform where five AI models (GPT, Claude, Gemini, Grok and Perplexity) work in one shared conversation, each reading what the others said before it answers. But the problem in this article is the one we work on every day: how to stop a fluent answer from reaching a person when it is wrong.

The lesson transfers directly. You do not make a generative system trustworthy by telling one model to be more careful. You make it trustworthy by creating independent chances for an error to be caught.

In Suprmind, the first chance is the models themselves. They share context, challenge each other’s claims and correct mistakes inside the thread. Across 1,324 real production turns in our Divergence Index, multi-model review surfaced at least one contradiction, correction or unique insight on 99.1% of them. But agreement is not proof. Five models can converge on the same wrong answer, which is why our AI Anti-Hallucinogen has a second, active layer. True North checks important claims, such as numbers, dates, named entities and citations, against external evidence, outside the models that wrote the answer. It currently runs in read-only shadow mode on a subset of eligible threads while we tune it.

A well-built voice agent follows the same logic with different parts. Retrieval and live tools supply the evidence. Deterministic checks guard the claims that carry risk. A person takes over where the evidence runs out.

Disagreement is information. Consensus is evidence. Neither is proof. On a customer call, proof is the order record, the payment confirmation and the policy actually in force. Build the agent so it can only say what those sources support.

FAQ

Can RAG stop voice AI hallucinations on its own?

RAG is the foundation, and a well-designed retrieval layer removes most invented answers about static knowledge such as policies and product details. It works best as one layer among several: live system queries for customer data, confirmation of spoken identifiers, validation around actions and checks on high-risk claims.

What is a hallucinated completion?

A voice agent telling the caller an action is done when it never happened, or when the backend returned an error. The fix is architectural: the agent may only confirm actions whose success response it holds, and the confirmation details come from that response.

Does lowering temperature reduce hallucinations?

It makes answers more consistent, not more correct. A model that misreads a policy at temperature zero repeats the same mistake on every call. Grounding, source routing and claim checks do the factual work.

Do voice agents have to tell callers they are AI?

In the EU, yes. Since 2 August 2026, Article 50 of the AI Act requires disclosure unless the AI nature is obvious, no later than the first interaction, which for a phone agent means the greeting. Other jurisdictions have their own rules, so check each market you serve.

Sources

author avatar
Radomir Basta CEO & Founder
Radomir Basta builds tools that turn messy thinking into clear decisions. He is the co founder and CEO of Four Dots, and he created Suprmind.ai, a multi AI decision validation platform where disagreement is the feature. Suprmind runs multiple frontier models in the same thread, keeps a shared Context Fabric, and fuses competing answers into a usable synthesis. He also builds SEO and marketing SaaS products including Base.me, Reportz.io, Dibz.me, and TheTrustmaker.com. Radomir lectures SEO in Belgrade, speaks at industry events, and writes about building products that actually ship.