The Voice on the Phone Was Not Your Mother

A familiar voice was the oldest identity check on earth — free, universal, automatic. AI withdrew it in about two years, and the FBI’s fix is a technology from the age of sieges: a secret family word.

At some point in the past couple of years, a new conversation has started happening at kitchen tables, and it would have made perfect sense to a Roman sentry. It goes like this: we need a word. A word nobody else knows — never written down, never texted, never dropped in the family group chat. If one of us ever calls in trouble, the other asks for the word. If the voice on the phone can’t produce it, hang up. Even if it sounds exactly like me. Especially then.

This is not advice from a prepper forum. It is the first protective tip in a public alert the FBI’s Internet Crime Complaint Center issued on December 3, 2024: “Create a secret word or phrase with your family to verify their identity.” The federal government’s recommended defense against the best voice forgery ever devised is, in other words, the authentication technology of a fort under siege.

The alert was specific about why. Criminals, it said, “generate short audio clips containing a loved one’s voice to impersonate a close relative in a crisis situation, asking for immediate financial assistance or demanding a ransom” — and, separately, “obtain access to bank accounts using AI-generated audio clips.” A month earlier, the US Financial Crimes Enforcement Network had issued its own alert, reported by CBC, warning that deepfake tools can “manufacture an apparent real event” and flagging the same family-emergency scheme. Nobody publishes a clean total for the damage; the numbers fold the AI version into the much older “grandparent scam.” But the Canadian Anti-Fraud Centre says Canadians reported losing nearly $3 million to emergency scams in 2024 — one country, one genre, and only what victims came forward about.

The raw material for the forgery is trivial. The Electronic Frontier Foundation says a few minutes of speech is enough to build a workable clone, and most of us have already donated that much to the internet: a voicemail greeting, a wedding toast, a podcast from 2019. The interesting half of the scam, though, is not the synthesis — it’s the reception. Recognizing a familiar voice isn’t a judgment; it happens upstream of judgment. You don’t authenticate your mother’s voice the way a bank inspects a signature. You just know it, and the knowledge arrives pre-loaded with everything else — relief, fear, the reflex to help. Fraud advice usually assumes a skeptical listener. This scam does its work in the half-second before skepticism exists, and the scripted crisis — the accident, the arrest, the ransom demand — is engineered to keep you there.

It is worth sitting with the speed of it. For the whole of human history, a known voice was the one credential that could not be counterfeited, the odd gifted mimic aside. In February 2023, the journalist Joseph Cox used a free clone of his own voice, plus his date of birth, to get into his Lloyds Bank account and scroll through his balances — past a system Lloyds marketed as “like your fingerprint,” analyzing “over 100 different characteristics.” Twenty-two months later, the FBI was telling families to invent passwords. That is roughly how long it took a security property our species got free at birth to go from load-bearing to recalled.

The biometric you broadcast

The institutional coda is quieter and worse. Banks and call centers spent the past decade building the voice into their security plumbing — TD, Chase and Wells Fargo run systems comparable to Lloyds’ — at precisely the moment the signal became forgeable. In November 2024, the BBC’s Shari Vahl passed voice-ID checks at Santander and Halifax with a clone cut from one of her own radio interviews, played through a basic iPad speaker. (She called from her registered number — a real attacker would also need the victim’s phone.) The banks’ responses have been serene. Santander said it had “not seen any fraud as a result of the use of voice ID,” which is a company grading its own homework, and Lloyds and Halifax called voice ID an “optional security measure” that still beats knowledge-based questions. Maybe so. But the fingerprint metaphor has a hole nobody at the banks seems to have noticed: a fingerprint stays on what you touch. A voice is left on everything you say in public, forever. They picked the one biometric we broadcast.

Rebuilt by hand

Underneath both stories — the panicked phone call and the bank login — is the same event: the withdrawal of a public good. Voice recognition was a security property that was free, universal and automatic, distributed at birth and maintained by evolution, with no enrollment, no hardware, no premium tier. It has been switched off, and nothing is replacing it at the same scale. The replacement is artisanal: one family at a time, agreeing on a word in a conversation somebody has to start, with relatives who must remember it under stress and never text it to anyone. The forgery scales industrially — one clip, a thousand calls. The defense is cottage industry. An AI developer named Asara Near floated the idea on Twitter in March 2023, calling it a “proof of humanity” word, and the phrase is more literal than it sounds: a token proving you are one of us. We had a technology for this once. It was the password — the sentry’s challenge, the shibboleth, the word you know rather than the body you are.

The easy irony is that the future of security is medieval. The less comfortable observation is that a piece of shared infrastructure has been withdrawn without replacement, and everyone finds out on their own schedule — some at the kitchen table, some on the phone with a voice that should not be trusted. The family password works precisely because it is nowhere: never posted, never synced, never sitting in a database waiting to be breached, a secret that exists in two heads and nowhere else. That is about the tightest definition of security left, and it propagates at the speed of dinner-table conversation, one household at a time, against forgeries that replicate like spam. The internet spent three decades making every utterance permanent, searchable and public. It turns out the only trustworthy utterance left is the one that never touched it.