Manav.id
Fraud ยท 17 min read

Synthetic identity fraud is a measurement problem before it is a fraud problem

A synthetic identity is built to be scored as real, and it usually is. The losses surface months later as credit charge offs with nobody available to call them fraud, which is why every published total is an estimate rather than a count. The control that works is not another signal at onboarding. It is a signature at the actions.

Picture a credit risk analyst at a mid size card issuer on an ordinary Tuesday, working through the monthly charge off file. One line stands out only because of how boring it is. Account opened fourteen months ago. Passed identity verification on the first attempt, no manual review. Thin file at origination, which is normal for the segment. Made small purchases and paid them in full for eleven months, never late, not once. Credit line increased twice on schedule because the behaviour justified it. Then, in the space of nine days, the line was drawn to the limit across four merchants, a balance transfer was requested, and the account went silent.

The analyst does the usual things. Calls the phone number on file, which rings and rings. Emails the address, which does not bounce and does not reply. Pulls the bureau file, which is intact and shows the same tidy history other lenders saw. There is no police report, because no consumer has called to say their identity was stolen. There is no dispute, because nobody disputes a charge they made themselves. There is nothing to investigate in the way an investigator would recognise, because every event in the account's life was genuine.

So the loss is coded as a credit loss. A borrower who looked good stopped paying. That is what the general ledger says, that is what the vintage analysis will absorb, and that is what the quarterly report will describe as normalisation in the thin file segment. The word fraud appears nowhere in the file.

Multiply that Tuesday across an industry and you arrive at the central strange fact about synthetic identity fraud, which is that the people who lose the most money to it are frequently unable to say how much they lost.

Can synthetic identity fraud be stopped without more risk scoring? Yes, but not at onboarding. A synthetic identity is engineered to be scored as real, so adding signals raises the attacker's cost slightly and the false positive rate for genuine customers considerably. The deterministic move is at the actions: bind the verified human to an enrolled device at account opening, then require that device's signature for credit line draws, payee additions and contact changes. Fifty synthetic accounts then require fifty live humans.

What is a synthetic identity, exactly?

Most coverage of this topic is vague in a way that makes the problem sound smaller than it is, so it is worth being precise.

A synthetic identity is a constructed persona assembled from a mixture of real and fabricated attributes. The typical construction anchors on one genuine identifier that belongs to a real person but is not actively used by them, then attaches invented or borrowed attributes around it: a name that does not match the identifier's true owner, a date of birth adjusted to suit the credit product, an address that receives mail, a phone number that answers, an email that has age on it.

The anchor matters more than anything else in the construction. It is what allows the persona to be found in the systems that lenders check, and it is chosen precisely because the true owner is unlikely to notice activity attached to it. Children, the recently deceased, people who have never borrowed, and people who do not monitor their credit are all overrepresented among anchors for exactly this reason.

Why this is not identity theft, and why the distinction matters

Classic identity theft impersonates a specific existing person in order to use that person's existing credit. It has a victim who notices, usually quickly, because the fraud shows up on their statement or their bureau file. The victim complains, the complaint is recorded, the loss is classified as fraud, and it gets counted.

A synthetic identity does not impersonate anyone. It creates somebody new. Nobody's existing credit is used, because the persona has no existing credit at the start. Nobody's statement shows an unexpected charge. The victim is the lender, and lenders discover their victimhood at charge off, if at all.

That single structural difference produces almost every downstream property of this fraud class, including the measurement problem, the reason detection struggles, and the reason the effective control sits somewhere unexpected.

Why does nobody know what synthetic identity fraud costs?

You will see large annual figures attached to this category. They circulate widely in vendor material, conference decks and trade press, usually in the tens of billions of dollars per year for the United States. Treat all of them as estimates constructed from assumptions rather than as counts, because that is what they are, and the honest version of this article has to say so before it says anything else.

Here is why a count is not available.

First, and most importantly, the loss is usually not classified as fraud in the first place. When an account with a clean payment history stops paying, the default classification is credit loss. Reclassifying it requires somebody to investigate and to conclude that the borrower never existed, which is expensive, retrospective and rarely worth doing for a single mid sized balance. Institutions that do this work systematically find synthetic identities in their charge off populations. Institutions that do not, do not, and both report their numbers to the same industry surveys.

Second, there is no agreed definition to count against. Different institutions, bureaus and regulators draw the boundary in different places, particularly around personas that mix real and fabricated elements in unusual proportions. The Federal Reserve has done substantial work over recent years attempting to establish a common definition precisely because the absence of one made the problem unquantifiable, and the existence of that work is itself the strongest evidence that the numbers predating it were not comparable.

Third, the surviving population is invisible by construction. Every published figure describes synthetics that were eventually identified. Synthetics that are still quietly paying their bills this month, still maturing, still eighteen months away from a bust out, are indistinguishable from good customers in every dataset anyone has. Any estimate is therefore a lower bound with an unknown gap above it.

None of this means the problem is small. The available evidence points the other way, and prosecutions regularly describe rings running large numbers of personas across many institutions. What it means is that anyone quoting a precise annual total to you, including a vendor with a chart, is quoting a modelled estimate. The correct posture is to treat the direction and the mechanism as well established and the magnitude as approximate.

For context on the wider measurement environment, the FBI's Internet Crime Complaint Center recorded reported losses of roughly 20.9 billion dollars across more than one million complaints in 2025 (FBI IC3 2025 Internet Crime Report). That figure is a count of complaints, and its greatest limitation is the same one that afflicts synthetic identity data: it can only contain what somebody reported. Synthetic identity fraud produces almost no complaints, which tells you where it sits relative to that total. We looked at which loss categories a signature would actually have prevented in the wedge fit re-read of the IC3 data, and the categories with no complaining victim are the hardest ones to reason about for exactly this reason.

What is bust-out fraud, step by step?

The operating pattern has a rhythm, and understanding the rhythm is what makes the control obvious later.

Construction. The persona is assembled and given whatever thin surface presence the target market expects: an address that accepts mail, a phone that answers, an email account with some age on it, sometimes a social presence.

Seeding. The persona applies for something easy. A secured card, a retail store account, a small credit builder product. Approval is not the point. The application itself creates a bureau record, and the record is what converts a persona with no history into a person with a file.

Cultivation. Small balances are carried and paid on time, month after month. This is genuine repayment with real money, and it is the part that defeats behavioural analytics, because the behaviour is not simulated. It is real behaviour, performed by an operator whose business model requires it. Cultivation typically runs for many months and sometimes for years.

Expansion. The now respectable file attracts pre approved offers and automatic line increases. The persona applies more widely, often across many institutions simultaneously, and each new approval further validates the file for the next lender. This is the compounding step, and it is why the same persona often appears at a dozen institutions before anyone loses money.

Bust out. Every line is drawn at once, frequently over a few days, sometimes with payments that will not clear timed to briefly inflate available credit. Then the persona stops responding and ceases to exist, which costs nothing, because it never existed.

Notice what is absent from that sequence: any moment where a system was deceived by a forgery. The bureau file is accurate. The payment history is accurate. The applications were completed correctly. The only false thing is the premise that a person is behind them, and no step in the process ever tests that premise directly.

Why do synthetic identities pass identity verification?

Because they are constructed to be scored well, and because scoring is what identity verification does.

Document checks can pass, since documents can be genuine or good enough, and since many thin file products do not require a document at all. Database checks can pass, since the anchor identifier is real and the persona's history is real. Velocity and device rules can pass, since a patient operator does not exhibit velocity. Behavioural analytics can pass, since the behaviour being analysed is authentic behaviour by a real operator who is genuinely paying bills.

This is the part that risk teams already know and that outsiders consistently miss: detection is not failing here because the models are poor. Detection is struggling because there is very little to detect. The signal that would separate this account from a genuine thin file customer does not exist in the data being examined, and no amount of model improvement conjures a signal that is not present.

There is a real analogy for this. Consider a librarian trying to spot a forged book among the shelves. If the forgery is a badly bound copy with the wrong typeface, examination works. If somebody has printed a genuine book, on genuine paper, with a genuine catalogue entry, and simply invented the author, then no amount of examining the book will help, because the book is not the thing that is false. You would have to go and look for the author.

That is the whole argument of this article in one image. The persona is not counterfeit. The person is missing.

Why will more signals not fix this?

Every additional signal added to an onboarding decision does two things, and institutions consistently plan for the first while absorbing the second.

It raises the attacker's cost, usually a little. A professional operator with an economic model and time will meet a new requirement, because the requirement applies to a persona that is being cultivated for months and a marginal cost per persona is amortised across a large expected draw.

It also raises the false positive rate for genuine customers, usually more than expected, and the genuine customers it catches are not randomly distributed. Thin file applicants, young people, recent immigrants, people who move often, people without a long address history, and people whose name is spelled inconsistently across databases are all more likely to look anomalous to a model trained on stable long file customers. These are frequently the applicants a lender most wants in a growth segment, and the cost of declining them is invisible on the fraud report and highly visible in the acquisition funnel.

So the marginal signal is a bad trade in both directions at once. It taxes the honest applicant more than the professional. This is the compounding pattern we have called detection debt: rising expenditure on classification against a classification rate that structural factors are pushing down.

What would a deterministic layer actually do?

Start from what is actually missing rather than from a product. What is missing is any link between the human who was verified at onboarding and the human performing the action that produces the loss. Identity verification is an event. An account is a duration. Nothing carries identity across the gap between them, which is the same discontinuity we described in the workforce context in identity continuity from hire to offboarding.

A deterministic layer closes that specific gap and nothing else. At onboarding, once the identity verification vendor returns a pass, the verified human enrols a device. The enrolment produces a key held on that device and a one way key derived from an on device face match, which is used to establish that a live human is present and that this human has not already enrolled elsewhere in the institution. No biometric template is stored anywhere. The face never leaves the phone.

Then, at a small number of designated actions, the institution requires a fresh signature from that device over the exact payload of the action. Not a session. Not a push notification that says approve. A signature over the specific bytes describing what is about to happen.

The arithmetic that makes this work

Here is a worked example. The numbers are illustrative and you should replace them with your own, but the shape is the point.

Suppose an operator wants to run fifty synthetic personas at one institution. Under a scoring regime, the marginal cost of persona fifty is close to the marginal cost of persona five: some data assembly, some cultivation capital, some patience. The work is largely parallelisable and largely automatable, which is precisely why professional operations exist at all.

Now add an enrolment requirement with liveness at onboarding and a signature requirement at the draw. Persona fifty now needs a fiftieth distinct live human, physically present, willing to complete an enrolment and then to be available again at the moment of the draw, months later. The operator's cost function changes from roughly constant per persona to roughly linear in recruited humans, and recruited humans are the most expensive, least reliable and most legally exposed input in any fraud operation.

That is the entire mechanism. It does not detect anything. It converts a data problem, which scales, into a human logistics problem, which does not.

What does the integration actually look like?

Deliberately small. The identity verification vendor keeps doing what it does. The risk engine keeps doing what it does. One binding is added at onboarding and one gate is added at the actions.

The binding record links an identity verification event to an enrolment without moving any of the underlying personal data:

{
  "type": "kyc_bound_enrollment",
  "idv_event_id": "prs_9f2c41ab",
  "idv_provider": "vendor-a",
  "idv_result": "pass",
  "idv_at": "2026-09-20T14:02:11Z",
  "enrollment_key": "ed25519:8bd1...c40f",
  "liveness": "passed",
  "uniqueness_scope": "issuer-consumer-cards",
  "biometric_stored": false
}

The gate at the action signs the action, not the session. A credit line draw payload looks like this, and the signature covers the hash of exactly these bytes:

{
  "action": "credit_line_draw",
  "account_id": "acct_5512",
  "amount": "9800.00",
  "currency": "USD",
  "destination": "ext_acct_7741",
  "requested_at": "2026-11-04T09:18:44Z",
  "nonce": "b17e9a2c"
}

And the server side check is unglamorous, which is the correct property for a control that has to run on every high value action:

receipt = manav.verify(signature, payload_hash)

if not receipt.valid:            reject("no human signature")
if receipt.key != account.enrolled_key:  reject("wrong human")
if receipt.age_seconds > 300:    reject("stale signature")
if receipt.action != "credit_line_draw": reject("wrong action")

approve()

The resulting receipt verifies offline against a published key, so a fraud investigator, an auditor or a regulator can check it later without calling the issuer's vendor and without any party having to be trusted at verification time. That property matters more than it sounds, and we go into it in the offline verification article.

Where should this sit in the risk stack?

Underneath it, and sold through it.

This deserves to be said plainly, because the alternative would be dishonest positioning. Identity risk is the most crowded category in the entire identity market. The established vendors in it are serious companies with large datasets, real science and years of tuning. Manav is not an identity verification vendor, does not run a bureau, does not score applicants, and would be a poor substitute for any of the incumbents at the job they actually do.

The argument is narrower and, we think, more useful because of it. Scoring decides whether to open the account. A signature decides whether the action executes. Those are different questions asked at different moments with different failure modes, and the second one currently has no answer at all in most institutions.

There is also a structural reason the deterministic layer wants to be neutral: a risk vendor cannot comfortably be the deterministic layer beneath its own competitor's score. Neutrality is a product feature here, not a positioning slogan.

ControlWhat it actually measuresWhat it cannot establish
Document and selfie verificationThat a document appears genuine and a face matched it at one momentThat the same human returns for later actions
Database and identifier checksThat the identifier exists and correlates with the claimed attributesThat a person stands behind the correlation
Velocity and device rulesThat this application resembles known bad patternsAnything about a patient operator with clean infrastructure
Behavioural analyticsHow the account behaves relative to a populationWhether authentic behaviour has an authentic owner
Bust out modelsThat an account now looks like it is about to bust outA control at the moment of the draw, rather than an alert about it
Enrolled human signature at the actionThat a specific live enrolled human authorised these exact bytesWhether that human is honest, coerced, or working for someone else

Read the last row carefully, including the right hand column. That is the honest boundary of this control and the next section is about it.

What this cannot do

Four limits, none of them small.

A recruited human defeats it. If an operator pays real people to enrol and to be available at the draw, every check passes, because a real human really is present and really is authorising. This is the money mule model applied to identity, it already exists at scale, and a presence proof does not close it. What the proof does is impose a cost per persona that scales with recruited humans and creates a durable link between a specific human and a specific fraudulent action, which is a materially better position for both loss prevention and prosecution than an unattributed charge off. It raises the floor. It does not end the category, and anybody claiming otherwise is selling.

First day fraud sits outside it. An account that is opened and drawn immediately never reaches a second action, so a control at the actions never fires. That case belongs to onboarding, which is where the existing risk stack is strongest, which is another reason this is a layer rather than a replacement.

Enrolment is itself a moment. Binding at onboarding inherits whatever assurance the onboarding had. If the identity verification that preceded the binding was defeated, the binding faithfully records a synthetic human's key. This matters more as capture pipelines come under attack, which we cover in the article on injection attacks against identity verification. Binding narrows the window from an account's entire lifetime to a single moment, and a single moment is a much better thing to defend, but it is not nothing.

Coverage is a rollout problem. Accounts opened before the control existed have no enrolled key, and retrofitting an existing book means asking established customers to enrol for reasons they did not ask about. The realistic sequence is new accounts first, then high exposure existing accounts at their next natural interaction, and a long tail that never gets covered.

What to do this week

  1. Pull a sample of last year's charge offs in your thin file segment and have somebody actually check whether the borrower existed. Most institutions have never done this, and the answer changes the size of the problem you think you have.
  2. Write down the list of actions that immediately precede loss in your product. Typically: first large draw, credit line increase acceptance, new external payee, address change, phone change, and card reissue to a new address.
  3. For each action on that list, write down what proof exists today that a human authorised it. In most stacks the honest answer for most rows is a session and a risk score.
  4. Measure your current false positive cost at onboarding, in declined genuine applicants, not in fraud caught. You need both numbers to evaluate any proposal, including this one.
  5. Ask your identity verification vendor what their event identifier looks like and whether it can be referenced by a downstream control without re-sharing the underlying personal data. That is the integration seam.
  6. Pick one action, the highest exposure one, and pilot a signature requirement on it for new accounts only. One action, one segment, measured against a control group.
  7. Instrument the pilot for both directions: loss prevented and genuine customers frustrated. A control that only reports the first number is not being evaluated, it is being marketed.
  8. Read the integration documentation and try the enrolment demo, which runs with no signup, so you can see what the customer experience actually costs in seconds before you argue about it internally.

Frequently asked questions

Can synthetic identity fraud be stopped without more risk scoring? Not entirely, but the largest loss class can be addressed at a different point. Scoring decides whether to open an account and will always be probabilistic. Requiring a signature from an enrolled device at the actions that precede loss makes those actions deterministic, and forces a fraud operator to supply one live human per persona rather than one dataset per persona.

Why do synthetic identities pass KYC checks? Because almost nothing in the persona is forged. The anchor identifier is genuine, the credit history is genuine because it was built with real payments, and the behaviour is genuine because a real operator produced it. Verification examines documents, databases and behaviour, and in this case all three are authentic. The only false element is the assumption that a person exists behind them.

What is bust-out fraud? It is the final stage of the synthetic identity lifecycle. After months or years of small balances repaid on time, which earns credit line increases and offers from additional lenders, every available line is drawn within a short window and the persona is abandoned. Because the account had a clean history and no complaining victim, the loss is usually recorded as a credit loss rather than fraud.

What is the difference between identity proofing and identity binding? Proofing establishes who somebody is at one moment, usually with documents and database checks. Binding attaches that verified human to a key on a device they hold, so later actions can be attributed to the same person. Proofing without binding means the account's whole subsequent life rests on an event that happened once and was never revisited.

Does this require storing customers' biometric data? No. The face match runs on the customer's own device and produces a one way key. No image and no template is transmitted or stored, which matters both for breach exposure and for the compliance profile under biometric privacy statutes. What the institution keeps is a public key and a set of receipts, neither of which is biometric data.

Will this increase onboarding abandonment? It adds one step at enrolment, so it will have some cost, and you should measure it rather than accept anybody's assurance including ours. The offsetting argument is that a deterministic control at the actions can allow a lower friction, less aggressive scoring posture at the front door, which is where abandonment is actually generated. Evaluate the two together, not separately.

Sources

  1. Federal Reserve, payment system fraud definitions and synthetic identity working material. federalreserve.gov/paymentsystems.htm
  2. FBI Internet Crime Complaint Center, 2025 Internet Crime Report, reporting roughly 20.9 billion dollars in reported losses across more than one million complaints. ic3.gov
  3. Social Security Administration, Electronic Consent Based Social Security Number Verification service documentation. ssa.gov
  4. Federal Financial Institutions Examination Council, guidance on authentication and access to financial institution services. ffiec.gov
  5. Financial Crimes Enforcement Network, advisories and analyses on identity related fraud typologies. fincen.gov/resources/advisories
  6. NIST Special Publication 800 63, Digital Identity Guidelines, on identity proofing assurance levels and their limits. csrc.nist.gov
  7. Consumer Financial Protection Bureau, materials on credit invisibility and thin file consumers. consumerfinance.gov
The persona is not counterfeit. The person is missing. You cannot find a missing person by examining the paperwork more carefully.