Your call center can't hear a deepfake. Stop asking it to.
Voice cloning did not break contact centre security by being clever. It broke it by removing the last thing that made a voice hard to forge. The answer is not a better detector, because there is no detector a human on a four minute handle time can run. The answer is to stop asking the agent to decide.
Picture a Tuesday afternoon on the floor of a mid sized retail bank's contact centre. An agent, eleven months into the job, picks up the two hundred and forty first call of the week. The caller is calm and slightly apologetic. He gives his date of birth, the last four of his card, and the name of the branch he opened the account at nineteen years ago. He explains that he has changed jobs, changed phone provider in the move, and needs the mobile number on the account updated before his new card arrives.
Every part of that is ordinary. The agent has done this several hundred times. She updates the number. Total call time, four minutes and ten seconds, which is under her target. She marks it resolved and takes the next call.
What she has actually done is hand over the account. The number she just wrote into the customer record is the number that will receive the one time passcode for the next login, the next card activation, the next payment confirmation. There was no malware, no breach, and no alert. The attacker did not break the bank's authentication. He used it, correctly, after persuading the one part of the system never designed to withstand a professional.
The voice was synthesised from a few dozen seconds of the real customer speaking, obtained the way anyone obtains it now: a podcast appearance, a voicemail greeting, a video on a relative's public profile. The clone did not need to be perfect. It needed to be unremarkable for four minutes to someone with no reason to be suspicious and no tool that would have helped if she had been.
You do not stop voice deepfakes in the contact centre by getting better at hearing them. You stop them by removing the agent from the authorisation decision. Sensitive changes, contact details, payout destinations, credential resets, should require a fresh cryptographic signature from the customer's enrolled device. A cloned voice can persuade an agent. It cannot produce a signature from a device it does not hold.
What actually happens when a cloned voice calls your contact centre?
It helps to be precise about the attack, because the popular version of this story is wrong in a way that leads people to buy the wrong thing.
The popular version is that a deepfake voice impersonates someone and steals money directly. That happens, and it is the version that gets written about. It is not the common shape, and not the one that should worry a contact centre leader most.
The common shape is quieter. The cloned voice is not used to authorise a payment. It is used to change an attribute of the account so that everything downstream falls over on its own: the email address, the mobile number, the postal address, the linked payout account, or a credential reset. That change is not itself a loss event and most fraud systems do not score it as one. It creates the loss hours or days later, through a channel that now looks entirely legitimate, because the account genuinely does have that phone number on it.
Why the quiet version is the dangerous one
Three reasons.
First, it does not trip anything. A payment out of pattern is what fraud monitoring exists to catch. A contact detail update by an authenticated caller during a support call sits inside the normal envelope of customer service activity, and many institutions do not generate a case for it at all.
Second, it defeats the controls layered behind it. If your step up for a large transfer is a passcode to the registered mobile, and the attacker has already changed the registered mobile, your step up now belongs to the attacker. This is the failure covered in the payee add is the real transaction: the industry hardens the visible high value action and leaves the configuration change that determines its outcome behind a weaker gate.
Third, it is cheap to retry. A failed attempt at moving money burns an account and often triggers investigation. A failed attempt at changing a phone number costs nothing. Call again in an hour and get a different agent.
The three doors into the contact centre
A cloned voice can be pointed at three different places, and they fail differently.
The IVR. Automated voice authentication, where a voiceprint is compared against an enrolled sample. This is the door most directly attacked by synthesis, because the comparison is entirely acoustic and no human is present to notice that the caller's story is odd.
The agent. Knowledge based authentication, where a human asks questions whose answers are supposed to be private. Voice cloning matters here not because the agent compares voiceprints but because a familiar sounding voice removes the cue that would otherwise prompt suspicion, and the knowledge questions have been answerable from breach data for a decade.
The escalation path. Supervisors, retention teams, and back office functions that resolve what the front line could not. More experienced people with broader permissions, frequently outside whatever authentication tooling the front line uses.
How much does contact centre voice fraud actually cost?
Here is where this post has to do something slightly unusual for a vendor, which is to tell you that most of the numbers in circulation on this topic, including numbers that appeared in an earlier version of this article, do not survive examination.
Why the numbers you have seen do not hold up
If you have read anything about contact centre deepfake fraud recently you will have met a small set of statistics that recur almost verbatim across vendor blogs, conference decks, and trade press: a total annual cost of call centre fraud, a percentage of leaders who rank voice deepfakes a top threat, a percentage who say they cannot detect them, a growth rate for synthetic voice in one sector.
We went looking for the primary publications behind those figures and could not establish them to a standard we are willing to publish. Several trace to secondary coverage citing other secondary coverage. Where a survey does sit behind a percentage, the sample is typically a few hundred self selected respondents recruited by a company that sells detection, which is a fine way to gather industry sentiment and not a defensible basis for a business case.
There is a deeper problem than sourcing, and it applies to every number in this category including any that Manav might one day publish. Consider what a vendor telemetry figure actually measures. A detection company reports that it observed some number of synthetic voice attempts across its customer base. That number counts the attacks the detector caught. It cannot count the attacks that succeeded, because a successful synthetic voice call is, by construction, one that nobody flagged as synthetic. The figure is therefore a floor whose distance from the ceiling is unknown and unknowable from that dataset.
Now consider a survey figure. When two thirds of respondents say they lack confidence in their ability to detect voice deepfakes, that is a genuinely interesting datum about industry sentiment. It is not a measurement of detection efficacy. It is a measurement of how detection efficacy feels to people who were willing to answer a survey about it, many of whom were recruited through channels that select for concern about the topic. We use that kind of finding in this article, and we label it as what it is.
This matters practically. If you take a headline loss figure into a budget meeting and someone in the room asks where it comes from, and the honest answer is that it comes from a vendor blog citing a trade publication citing an unnamed study, you have damaged your own case. The mechanism argument is stronger than the statistics argument and does not have that failure mode.
What the evidence does support
Several things are documented well enough to build on.
Voice cloning from short samples is real, cheap, and available to non specialists. The United States Federal Trade Commission ran a public Voice Cloning Challenge to solicit approaches for preventing harms from the technology, which is a regulator formally acknowledging that the capability is widespread and the defensive answer is not obvious.
Impersonation to obtain account access is a large, tracked category of crime. The FBI's Internet Crime Complaint Center publishes annual reporting on account takeover and issues public service announcements on specific techniques. These are complaint driven figures and therefore floors, but floors established by a law enforcement agency rather than by a party with a product to sell. Separately, NIST guidance on digital identity has for years discouraged reliance on factors that can be replayed or captured without the subject's cooperation.
That is enough. You do not need a contested dollar figure to justify a control whose absence is demonstrable from first principles.
Why can't voice biometrics catch a deepfake?
Let us be fair to voice biometrics before criticising it, because it solved a real problem and it is still solving part of it.
Before voiceprints, contact centre authentication was knowledge based: mother's maiden name, first school, last transaction amount. That model was already broken by the mid 2010s, because the answers sat in breach corpora and could be bought. Voice biometrics improved on it substantially. It authenticated something the caller was rather than something they knew, it worked passively, and it cut handle time. Institutions that deployed it got a genuine reduction in fraud and a genuine improvement in customer experience at once, which is rare.
The property that made it work is the property that broke it
Voice biometrics works by comparing an incoming acoustic sample to an enrolled reference and producing a similarity score. That is a detector. It answers the question "does this sound sufficiently like the enrolled person" with a probability.
Every detector in security lives or dies on one question: can the adversary iterate against it? If the adversary gets unlimited attempts and can tell whether each worked, the detector's advantage decays toward zero, because the adversary runs an optimisation loop against a target that only moves when the vendor ships an update.
Voice is the ideal case for the adversary. The reference material is public. Generation models improve monthly and are sold commercially. Attempts are free and feedback is immediate, because the call either proceeds or it does not. And unlike a face presented to a controlled sensor, a voice arrives over a telephone channel that has already destroyed most of the acoustic detail a detector might have used.
An analogy that gets the shape right. Imagine you hire a doorman and tell him to admit anyone whose handwriting matches a signature card in his drawer. That works while forging handwriting is a skill. It stops working the day a machine can produce a perfect match from a photograph of the signature, and it stops working completely once that machine is free and the forger can try again tomorrow with a different doorman. You do not fix that by hiring a doorman with better eyesight. You fix it by no longer using handwriting as the thing that opens the door.
The handle time problem nobody costs properly
There is a second failure mode that has nothing to do with acoustics.
Contact centre agents are measured on average handle time, first call resolution, and customer satisfaction. Those are the right things to measure for the job as designed. None of them reward scepticism. An agent who challenges a caller and turns out to be wrong has created a bad customer experience, a complaint, and possibly a coaching conversation with their supervisor. An agent who processes a fraudulent request has done what the process told them to do and will very likely never learn that anything happened.
So the incentive gradient runs toward compliance with the caller. We are asking that person to out interrogate someone who does this professionally, has rehearsed, and has the customer's real data in front of them. This is not a training problem. Training is what organisations reach for because it is cheap and legible, and it produces a measurable improvement in a simulated test and very little against a real adversary who adapts. The same structural point is made at greater length in detection debt.
Would caller ID authentication help?
Worth answering properly, because many people assume it is already handled. Caller ID authentication frameworks let an originating carrier cryptographically sign the calling number and assert how much it knows about the caller's right to use it, and let the terminating carrier verify that signature. In the United States this is deployed under regulatory mandate and it has meaningfully reduced the crudest forms of number spoofing.
It does not help here, for a reason that is precise rather than dismissive. The framework attests to the provenance of the number, not to the identity of the person speaking. An attacker who lawfully obtains a number from a compliant provider gets a fully attested call every time. And nothing in the framework speaks to call content, so a perfectly attested call carrying a cloned voice verifies exactly as well as any other. We wrote that argument up separately in caller ID authentication proves the carrier.
This is the recurring pattern across this whole area. Each layer authenticates the thing it was designed to authenticate, correctly, and the fraud moves to the layer nobody is attesting.
How do you stop it without a better detector?
Change what the call is for.
Today a support call is both a conversation and an authorisation channel. The agent listens, forms a judgement, and executes a change. The authorisation is the agent's belief.
Instead, let the call stay a conversation and move the authorisation somewhere the caller cannot reach. When a sensitive change is requested the agent does not execute it, the agent initiates it. The system sends a signing request to the customer's enrolled device, the customer approves it there on hardware the attacker does not hold, and the change completes only when a valid signature returns.
The attacker can be word perfect. He can know the branch, the date of birth, the last four transactions, and he does sound exactly like the customer. None of it produces the signature, because the signature requires possession of a specific device and a local unlock on it. Persuasion does not enter into it.
What is actually signed
The important detail, and the one that separates this from a push notification that says "approve login", is that the signature covers the specific change, not the session. The customer is not approving "some activity is happening on your account". They are approving this exact modification, rendered on their own screen, on a device the caller does not control.
A canonical payload for a mobile number change looks roughly like this. The exact field names do not matter; the properties do.
{
"action": "contact.mobile.update",
"account_ref": "acct_9f2c41",
"current_value": "+44 7700 900311",
"new_value": "+44 7700 900884",
"requested_via": "voice_support",
"agent_ref": "cc_agent_4471",
"case_id": "CS-2026-118842",
"issued_at": "2026-10-02T14:11:03Z",
"expires_at": "2026-10-02T14:16:03Z"
}
That object is canonicalised to a deterministic byte sequence, hashed, and the hash becomes the challenge for a WebAuthn assertion on the customer's enrolled device. What comes back is verified like this.
payload = canonicalise(change_request) # deterministic bytes
challenge = sha256(payload)
assertion = await customerDevice.sign(challenge) # passkey, local unlock
ok = verify(
assertion,
challenge = challenge,
credential_id = enrolled_credential_for(account_ref),
origin = "https://bank.example",
max_age = 300 # seconds
)
if not ok: reject("no valid customer signature")
apply(change_request)
store_receipt(assertion, payload) # offline verifiable later
Four properties do the work. The signature covers the payload, so if any field changes the hash changes and the assertion no longer verifies. The key lives on the customer's device and cannot be exported, so possession is not transferable by persuasion. It expires in minutes, so a captured signature is worthless later.
The receipt survives. What you keep is not a log entry your own system wrote about itself. It is an artifact signed by the customer's key, verifiable offline against a published key by a dispute team, an auditor, or a regulator who trusts neither party. That distinction matters more than it first appears, and we develop it in the piece on why callback verification fails.
Where Manav sits, precisely
To be exact about what is shipped rather than aspirational: the per action passkey signature bound to a payload hash, the offline verifiable receipt, and the companion device pairing flow with on device face match and liveness are shipped and demonstrable at the signing lab. Native connectors into specific contact centre platforms are not; integration today is through the API and the widget, documented in the developer docs. We would rather you knew that before a procurement conversation than after one.
Which changes should require a signature?
Not all of them, or you will have built an unusable contact centre and the business will route around you within a quarter.
The useful sorting question is not "how sensitive does this feel" but "if this change is fraudulent, what does it enable, and can we undo it". Rank by what the change unlocks downstream rather than by its apparent value.
| Requested change | What it unlocks if fraudulent | Reversible? | Gate |
|---|---|---|---|
| Mobile number or email on file | All future one time passcodes and recovery | Yes, but usually after loss | Signature |
| Payout or linked bank account | Direct movement of funds | Rarely | Signature |
| Credential or authenticator reset | Durable independent access | Yes, if noticed | Signature |
| Postal address for card despatch | Physical card in attacker's hands | Partially | Signature |
| Adding an authorised third party | Ongoing legitimate looking access | Yes | Signature |
| Raising a transaction limit | Larger single loss | Yes | Signature above threshold |
| Statement copy, balance enquiry | Information only | N/A | None |
| Disputing a transaction | Nothing directly | Yes | None |
| Card freeze or block | Nothing, it is protective | Yes | None, never gate safety actions |
That last row deserves emphasis. Never put friction in front of an action whose effect is to reduce risk. If a customer calls to freeze a card, freeze the card. The failure mode of a fraudulent freeze is an inconvenienced customer. The failure mode of a delayed freeze is a drained account.
What does the agent's day look like afterwards?
Better, which is the part that surprises people who expect this to be a friction story.
Today, when a call feels wrong, the agent has two bad options. Proceed and risk processing fraud, or refuse and create a complaint from a customer who is probably legitimate, with no evidence behind the refusal. Most proceed, and those who refuse are frequently overruled on escalation.
With the change gated on a signature there is a third option that beats both: initiate the request and let the customer approve it. The agent accuses nobody. The system asks for confirmation from the device already associated with the account. A genuine caller taps and the call continues. A fraudulent one produces an excuse, and the excuse is itself the signal the agent could not previously obtain.
It also removes a burden the industry has quietly placed on frontline staff for years: asking people evaluated on speed to be the final control preventing six figure losses, then treating it as a personal failure when they lose to a professional. Moving the decision into cryptography is, among other things, a better deal for the agent.
The same reasoning applies with even more force where the insider is the target rather than the caller. When a support agent can be bribed rather than merely deceived, no amount of training helps, and the only control that survives is one the agent does not hold. We work that through in when support staff can be bought.
Honest limits
This is not a complete answer, and the places it falls short are worth stating plainly.
It only protects enrolled customers. Enrolment takes time and never reaches one hundred percent. You will run a mixed estate for years, and during that period the unenrolled population is protected by whatever you had before. Prioritise enrolment by account value and by the customers whose details have already been targeted.
It moves the attack to enrolment. If an attacker can enrol their own device against someone else's account, they have bypassed everything downstream. Enrolment is now the highest value moment in the lifecycle and needs to be treated that way, which is a different problem from the one this article solves.
A compromised customer device defeats it. If malware on the phone can approve prompts, the signature proves possession of a compromised device. That raises the attacker's cost from a phone call to device compromise, which is a large increase and not an infinite one.
It does not address first party fraud. A customer who authorises a change and later disputes it has produced a valid signature. The receipt helps with the evidentiary question afterwards, since you hold an artifact the customer's own key produced rather than your account of what happened.
Accessibility is a real constraint, not a footnote. Some customers have no smartphone, some have cognitive or motor impairments that make device interaction hard, and some are calling precisely because their device is lost. Every one of those people needs a path that works, and that path must be staffed and treated as a normal route rather than an exception grudgingly granted. A control that locks out vulnerable customers to stop fraudsters has failed at the thing the institution actually exists to do.
It does not replace your fraud operation. Monitoring, case management, and investigation all still matter. This removes one class of loss deterministically. It does not remove the rest.
What to do this week
- Enumerate every change an agent can make without the customer touching anything. Include supervisor and back office permissions, and include the systems that are not the primary CRM. Most institutions are surprised by the length of this list.
- Sort that list by what each change unlocks downstream, not by how sensitive it feels. Anything that redirects a communication channel, a payment, or a credential goes at the top regardless of how routine it seems.
- Measure your exposure window. Pull twelve months of contact detail changes and calculate the median time between a change and the next authenticated action using it. That is how long an attacker has.
- Check whether a changed detail is usable immediately. Many institutions allow a newly added number to receive a one time passcode straight away. A short cooling period on newly changed contact channels costs almost nothing and removes the fastest version of this attack today, before any project starts.
- Instrument the escalation path. Confirm supervisor overrides and retention team actions are logged against an individual rather than a shared account. If your audit trail names a team, you have an attribution problem before you have an authorisation problem.
- Pick one change type and gate it. Payout destination is usually the cleanest first candidate: low volume, high consequence, and no reasonable customer objects to confirming it. Prove the flow there before touching anything high volume.
- Write down your fallback path before you launch, and staff it. Decide in advance what happens for the customer with no device, and make sure that route is not simply "the agent uses judgement", which is the control you were trying to remove.
- Stop citing unsourced statistics internally. Rebuild your business case on exposure arithmetic from your own data, which nobody can dispute, rather than on a number from a vendor deck, which somebody will.
Frequently asked questions
Is voice biometrics now useless? No. It remains useful as a low friction signal for routine, low consequence interactions, and it still raises the cost of casual impersonation. What it cannot do is carry the authorisation decision for a consequential change, because it is a probabilistic comparison against material the attacker can obtain and iterate against freely.
Does this add friction to every call? No, and it should not. Routine enquiries, balance checks, disputes, and anything protective such as freezing a card should be untouched. The signature applies only to the small set of changes that redirect a communication channel, a payment, or a credential. In most contact centres that is a low single digit percentage of call volume.
What if the customer has no smartphone or has lost their device? They take a documented fallback path, which might be a second enrolled authenticator, a hardware key, or an in person or staffed high assurance route. The important design rule is that the fallback must be resourced and treated as normal, and it must not resolve to an agent making a judgement call, since that reintroduces exactly what you removed.
How is this different from sending a one time passcode? A passcode proves that whoever is holding the receiving channel right now received a message, and that channel is often the very thing the attacker is trying to change. A signature is produced by a key that cannot be exported from the enrolled device, and it covers the specific change rather than the session, so it cannot be relayed or reused.
Can we just buy deepfake detection for the audio stream? You can, and it may be worth something for triage and prioritisation. Understand what you are buying: a probability, produced by a model, against an adversary whose generation model updates more often than your detector and who gets unlimited free attempts with immediate feedback. Treat it as a signal that helps you look in the right place, not as the control that decides.
Where do we start if we have thousands of agents and legacy systems? One change type, one queue, one region. Payout destination changes are usually the best first target because volume is low, consequence is high, and customer acceptance is easy to obtain. Prove the enrolment path and the fallback path there, measure the completion rate honestly, and expand only once both work.
Sources
- Federal Trade Commission, Voice Cloning Challenge, a public competition on preventing harms from voice cloning technology: ftc.gov
- FBI Internet Crime Complaint Center, annual Internet Crime Reports, including account takeover and impersonation categories: ic3.gov
- FBI Internet Crime Complaint Center, public service announcements on specific fraud techniques: ic3.gov/PSA
- NIST Special Publication 800-63B, Digital Identity Guidelines, on authenticator classes and restricted authenticators: nist.gov
- Federal Communications Commission, call authentication and caller ID framework: fcc.gov
- W3C Web Authentication, the specification underlying passkey assertions: w3.org
- Cybersecurity and Infrastructure Security Agency, advisories on social engineering of support functions: cisa.gov
A cloned voice can convince your agent. It cannot hold your customer's phone. Put the decision where the phone is.
This lesson is part of Proof of human intent, the guide to the whole problem area.