Manav.id
Regulated ยท 17 min read

The same volunteer, three trials, two identities.

A participant enrolled at two sites for two studies is not primarily a data quality problem. They are a person receiving two investigational products that nobody assessed together. The industry answers this with shared registries of identifying information, which is a bigger database of exactly the thing participants are most reluctant to hand over.

Picture a research coordinator at a mid sized site on a Thursday afternoon. She is screening for a phase II study in an inflammatory indication. The participant in front of her is punctual, articulate, and answers the eligibility questions with unusual fluency. He knows what a washout period is. He uses the word titration correctly. He has clearly done this before, which is not disqualifying, because prior trial participation is normal and often desirable.

What she cannot know, working from the file in front of her, is that he completed screening for a competing study eleven days ago at a site forty miles away, using a slightly different spelling of his surname and a date of birth off by one digit. He disclosed no active enrolment on either form. Both sites will dose him. Both sponsors will collect his data. Neither will know that the pharmacokinetic profile they are building describes a person carrying another compound.

If something goes wrong, the adverse event will be attributed to whichever drug the reporting site is studying, and the causality assessment will be made by an investigator working from a medical history that omits the most important fact in it.

Short answer: Duplicate enrolment is a uniqueness of human problem, and it is dangerous before it is expensive. Sites cannot detect it because they match on identifiers a determined participant simply varies. Shared registries work but require sites to submit personal data to a third party. The better shape is a device held one way key that answers "is this person enrolled elsewhere in this window" without anyone learning who they are.

What is a professional subject, and what is a duplicate subject?

The terms are used loosely, so it is worth separating them, because they describe different behaviours with different remedies.

A duplicate subject is someone participating in more than one study at the same time, or re-entering the same study after being screened out. Registry operators in this field, including CTSdatabase, define a duplicate subject as a participant identified as taking part in another trial or as having very recently completed one. The overlap is the harm.

A professional subject is someone who participates repeatedly and habitually, travelling between sites, and who may omit or alter information in order to remain eligible, with the collection of multiple stipends as a motivation. CTSdatabase describes this population in similar terms. Repeat participation on its own is not misconduct. Concealment to defeat eligibility criteria is.

Both groups exist. Both have existed since paid clinical research has existed. Vendors operating in this space, including Verified Clinical Trials and CTSdatabase, describe duplicate and professional subjects as structural risks in modern research rather than isolated anomalies, which is a claim from parties who sell the remedy and should be read with that in mind, though it is also consistent with what site staff report.

Why does duplicate enrolment matter?

Three distinct harms get bundled together in most discussions of this topic, which muddles the argument. They are not equally important and they do not have the same remedy, so separate them.

Participant safety comes first

This is the one that matters most and gets discussed least, probably because the other two are easier to put in a business case.

Eligibility criteria exclude concurrent investigational products for a reason. Two compounds in one body produce interactions that nobody has studied, because the entire purpose of the trial is that the interactions of even one compound are not yet fully known. Washout periods exist because residual drug confounds both safety and efficacy. A participant who defeats these criteria is not gaming a paperwork requirement. They are removing the protection that the protocol was designed to give them, and they are doing it without the information needed to understand the risk they are accepting.

The investigator's ability to assess causality also degrades. When an adverse event occurs, the investigator asks whether the study drug is a plausible cause given everything known about the participant. If the most relevant fact is missing, that assessment is being made on a false record, and the resulting safety signal, or absence of one, propagates into a regulatory submission.

Data integrity comes second

A duplicate participant contributes contaminated data to both studies. In a large phase III trial, one participant is unlikely to move an endpoint. In an early phase study with a few dozen participants, or in an indication where a small effect size is being chased, a handful of contaminated records can matter, and they are indistinguishable from noise because there is nothing in the data marking them as different.

The insidious property is that this contamination is invisible in the usual quality checks. The data is internally consistent. The visits happened. The samples are real. Nothing looks wrong, which is a theme that recurs in every failure this blog covers.

Cost comes third

Screening, dosing, monitoring and following a participant is expensive, and a fraudulent participant consumes an enrolment slot that a legitimate one could have filled, which delays completion. Clinical development costs are large and widely debated, with published estimates varying enormously depending on methodology, what is counted, and whether failed programmes are amortised in, so we are not going to attach a number to a single duplicate participant. It is enough to say that enrolment slots are scarce and expensive and that delay is the dominant cost driver in most programmes.

HarmWho bears itDetected byFixable after the fact?
Concurrent investigational productsThe participantUsually nothing, unless an event occursNo
Corrupted causality assessmentFuture patients, the sponsor, regulatorsRarely, and only retrospectivelyNo
Contaminated efficacy dataThe sponsorStatistical anomaly detection, sometimesPartially, by exclusion, if identified
Wasted enrolment slot and spendThe sponsor and the siteOnly if the duplicate is foundNo, the time is gone
Delay to programme completionThe sponsor and future patientsAggregate schedule slippageNo

Why do participants enrol twice?

It would be easy and wrong to write this section as though the participants were simply dishonest. The incentive structure deserves an honest look.

Compensation for participation is legitimate and necessary. Trials ask people to give up time, travel to sites, undergo procedures, and accept risk. Paying for that is standard practice and is permitted by regulators, with institutional review boards scrutinising payment levels precisely because payment that is too high becomes undue inducement. Ethics bodies have argued about where that line sits for decades and there is no settled answer.

Now consider a person for whom trial stipends are a meaningful fraction of income. The system has told them, correctly, that their participation has value and deserves payment. It has also constructed eligibility criteria that they experience as arbitrary gatekeeping between them and that payment. It should not be surprising that a small number of people optimise against those criteria. Most participants are altruistic, many participate because they or someone close to them has the condition being studied, and the professional subject population is small. But it is not zero, and the incentives are not accidental.

Decentralised and remote trial designs, which expanded substantially from 2020 onward and which regulators including the FDA have issued guidance to support, made participation easier for everyone. That is a genuine good: they widened access for people who cannot travel to an academic medical centre. They also removed the in person friction that used to limit how many sites one person could realistically reach, and coverage in 2025 noted the resulting rise in concerns about fraudulent participants in health research. Both things are true at once.

How do sponsors check for duplicates today?

The current answer is subscription registries. A site submits participant identifiers to a shared database at screening, the database checks them against records submitted by other subscribing sites, and returns a match or no match. Verified Clinical Trials and CTSdatabase are the established operators in this space. Reporting has described an insurer, Chubb, partnering with Verified Clinical Trials to address duplicate enrolment before it compromises safety or skews results, which is worth noting for a specific reason: when an insurer partners on a control, the risk is being priced, and priced risk is real risk.

These registries work. They are the sensible response given available technology, and sponsors who use them find duplicates that would otherwise have gone undetected.

They also have two structural properties worth examining.

They require identifying data to leave the site

To match a participant across sites, the registry must receive something identifying: name components, date of birth, sometimes partial identifiers or biometric derived values. That means a database exists containing the fact that a specific named person participated in research on a specific date, aggregated across sponsors.

Consider what that record is. Participation in a trial reveals, by inference, that a person has or is at risk of a specific condition. A registry of research participants is therefore adjacent to a registry of who has which diseases, held by a private party, outside the health system, for an operational purpose. It is handled carefully by the operators and it is still the single most sensitive category of data most participants will ever be asked about, and it must be created before the protective benefit is delivered. Institutional review boards and participants notice this, which limits adoption, which limits the network coverage that makes the registry useful.

They match on things a determined participant can vary

Identifier matching is probabilistic. A professional subject who wishes to avoid detection can vary a middle initial, transpose two digits of a date of birth, use a different address, or present a different phone number. Matching algorithms compensate with fuzzy logic, which raises false positives, which means legitimate participants get flagged and site staff spend time resolving them.

This is the same shape as the duplicate detection problem elsewhere, and we go through the general version, including the real harm of false merges, in one human, many accounts, one account, many humans. The registry is doing identity resolution, and identity resolution is hard precisely when the subject is motivated to defeat it.

What would uniqueness without a registry look like?

The question a site actually needs answered is narrow: is this person already enrolled in another study right now? Notice that answering it correctly does not require knowing who the person is. The site already knows who they are. Nobody else needs to.

That distinction is the whole design. Identity resolution asks "who is this", which requires identifying data. Uniqueness asks "is this the same one as that", which does not.

Here is the shipped, near term version, and we are going to be precise about what is available today versus what is proposed, because overstating this would be exactly the kind of claim this blog exists to criticise.

At the consent visit, the participant enrols a key on their own device. The enrolment produces a one way key derived on the device: no image is transmitted, no biometric template is stored anywhere, and the derived value cannot be reversed into a face or a name. The participant then signs an enrolment attestation, and the resulting receipt goes into a wallet they control.

{
  "action": "trial.enrollment",
  "protocol_id": "SPN-2026-0114",
  "indication_class": "INFLAM-2",
  "site_id": "SITE-0442",
  "enrolled_at": "2026-09-25T14:02:11Z",
  "window_ends": "2027-03-25T00:00:00Z",
  "subject_key": "sha256:8b41...e7d2",
  "attestation": "no other active investigational enrollment"
}

The subject key is derived on the participant's device and is stable for that person. The receipt is signed by the participant and countersigned by the site. It contains no name, no date of birth, and no address. A second site, at a second screening visit, can ask the participant to present receipts and verify them offline:

const keys = await fetchOnce('https://manav.id/.well-known/manav-keys')

for (const r of presentedReceipts) {
  if (!verifyEd25519(keys[r.kid], canonicalJson(r.payload), r.signature)) continue
  if (r.payload.subject_key !== thisSubjectKey) continue
  if (Date.parse(r.payload.window_ends) > Date.now()) {
    flag('active enrollment elsewhere', r.payload.protocol_id)
  }
}

Verification needs no call to us and no call to the other sponsor. It checks a signature against a published key. If the participant is the same human, the subject key matches, and the overlapping window is visible without either sponsor learning anything about the other's participant list beyond what the participant chose to present.

Now the honest part. This design is consent based: it works when the participant presents their receipts. A participant determined to conceal a concurrent enrolment can decline to present, or enrol a second key on a second device. Closing that gap requires a window scoped uniqueness proof, a nullifier, that lets a sponsor learn that some valid participant has already enrolled in this indication window without learning which one, and without the participant being able to generate a second unlinked identity for the same window. That construction exists in the research literature and in deployed systems such as Semaphore, and it is a roadmap item for us, not a shipped feature. We are describing where this goes, not what you can buy on Monday.

Why is a bigger registry the wrong direction?

The instinctive industry response to incomplete coverage is to grow the registry. More sites, more sponsors, more identifiers, better matching.

That direction has a problem that gets worse rather than better with scale. A registry's usefulness grows with coverage, and so does its sensitivity. A complete registry of research participants, aggregated across sponsors and indications, would be one of the most sensitive private health databases in existence: a lookup from a person to the categories of disease they have been studied for. Its protective value would be maximal at exactly the moment its breach consequences became severe.

It is also a consent problem that compounds. Every expansion requires more participants to agree that their participation may be disclosed to a third party, and the participants most likely to refuse are the ones most concerned about privacy, who are not the same population as those enrolling twice. You lose coverage where you least want to lose it.

The alternative is not less protection. It is protection that does not accumulate a liability, which is the argument we make generally in behavioural biometrics is surveillance with better branding. A per human one way key can establish uniqueness within a window while leaving no database of who participates in what.

What this cannot do

It says nothing about eligibility misrepresentation. This is the big one and it deserves to be first. A participant who lies about their smoking history, exaggerates symptom severity to qualify, or conceals a comorbidity is committing the other major integrity failure in clinical research, and uniqueness proofs are entirely silent about it. Some evidence suggests misrepresentation is the more common problem. Nothing here touches it.

It requires adoption to be useful. A uniqueness check across sponsors only works if sponsors participate, which is the same network problem the registries have. We would be replacing one coverage problem with another, and we should say so rather than pretend a new mechanism starts with full coverage.

Consent based presentation is defeatable. Stated above and worth restating. Until window scoped uniqueness proofs ship, a participant who declines to present receipts is where we started.

Device enrolment excludes some participants. Trial populations include elderly participants, people who are acutely unwell, and people without smartphones. Any design that assumes a personal device excludes exactly the populations research most needs to include, so a site administered fallback is mandatory, not optional.

We have no clinical systems integration. Manav does not integrate with electronic data capture or eConsent platforms today. Anything described here is a design and a pilot shape, not a deployment.

What to do this week

  1. Separate the three harms in your own risk register. If duplicate enrolment is filed only under data quality at your organisation, the safety case is not being made to the people who fund controls.
  2. Ask your registry vendor for their false positive rate and the coordinator time spent resolving flags. That number determines whether site staff take the check seriously or route around it.
  3. Look at what identifying data your sites currently transmit to third parties at screening, and check that your consent language actually covers it in the way participants would understand.
  4. For decentralised arms specifically, write down what identity assurance you have at enrolment. In many remote designs the honest answer is a document photograph, and document photographs are no longer strong evidence, for reasons covered in the camera is no longer evidence.
  5. Ask whether your eConsent platform can attach a participant signed receipt to the consent record. Several can produce signature artifacts already and the capability is unused.
  6. If you are a sponsor with an insurer engaged on trial risk, ask what evidence of enrolment integrity would affect pricing. That conversation moves budgets faster than any internal argument.

You can see the underlying signing and verification flow in the walk up demo, and the integration surface is described in the developer docs.

Frequently asked questions

How do you prevent duplicate enrolment in clinical trials without a shared database of participants? Bind each participant to a device held one way key at the consent visit, then have them sign an enrolment receipt naming the protocol, the indication class and the exclusion window. A second site verifies the receipt offline against a published key. The check establishes that the same human is already enrolled without any registry learning who that human is.

What is a professional subject in clinical research? A participant who takes part in studies repeatedly and habitually, often travelling between sites, and who may omit or alter eligibility information in order to remain qualified and collect multiple stipends. Registry operators including CTSdatabase define the category in these terms. Repeat participation by itself is normal and often valuable. Concealment to defeat eligibility criteria is the problem.

Why is duplicate enrolment a safety issue and not just a data issue? Because exclusion criteria and washout periods exist to prevent participants receiving investigational products whose interactions have never been studied. A duplicate participant removes that protection from themselves without the information needed to weigh the risk, and the investigator assessing any resulting adverse event does so from a medical history missing its most relevant fact.

Does the FDA require identity verification for trial participants? Regulators require investigators to confirm eligibility and to maintain accurate records under good clinical practice, and electronic records and signatures are governed in the United States by 21 CFR Part 11. There is no single federal mandate prescribing a specific identity verification technology for participants, which is why practice varies widely between sites and sponsors.

Do shared subject registries actually work? Yes, within their coverage. Sites that subscribe find duplicates they would not otherwise catch. The limits are that matching relies on identifiers a motivated participant can vary, that coverage depends on how many sites subscribe, and that the model requires identifying data about research participation to be held by a third party, which some review boards and participants resist.

What is a nullifier and why does it matter here? A nullifier is a cryptographic value that lets someone prove they have not already acted within a defined scope, such as one enrolment per indication per window, without revealing their identity or linking their actions to each other. Deployed systems such as Semaphore implement this pattern. It would allow uniqueness checks that do not depend on the participant voluntarily presenting receipts. It is a roadmap item for us rather than a shipped capability.

Would this replace the registries? Not initially and possibly not ever. A privacy preserving uniqueness layer and an identifier registry answer overlapping questions with different tradeoffs, and sponsors under regulatory scrutiny will reasonably want more than one control. The realistic path is composition, where receipts reduce how much identifying data has to move for the same protective effect.

Sources

  1. CTSdatabase, definitions of duplicate subject and professional subject in clinical research: ctsdatabase.com
  2. Verified Clinical Trials, on duplicate enrolment and professional subjects as structural risks, and its reported insurer partnership: verifiedclinicaltrials.com
  3. US Food and Drug Administration, guidance documents including those covering decentralised clinical trial conduct: fda.gov
  4. 21 CFR Part 11, Electronic Records and Electronic Signatures: ecfr.gov
  5. International Council for Harmonisation, efficacy guidelines including good clinical practice: ich.org
  6. US Department of Health and Human Services, Office for Human Research Protections, on participant payment and undue inducement: hhs.gov/ohrp
  7. Semaphore, a deployed implementation of nullifier based uniqueness proofs, as reference for the privacy preserving construction described: semaphore.pse.dev

Note on figures: we have not attached a cost per duplicate participant or a prevalence rate. Published clinical development cost estimates vary by an order of magnitude depending on methodology, and the duplicate rates that circulate come from vendors whose measurements are limited to their own subscriber base and whose detected rate is by definition a floor rather than a count. Both would have been easy to quote and neither would have been honest.

The registry has to learn who you are before it can protect you from being enrolled twice. That trade was never necessary.