Manav.id
Research ยท 18 min read

A third of your survey responses may not be from people

Industry estimates of fraudulent online survey responses run from 15 percent to 40 percent depending on who is counting and how. The spread itself is the finding: nobody can measure this well, because every measurement is a detector, and the things being detected now write better open ends than your respondents do.

The Tuesday the data looked perfect

Picture a researcher at a mid-size agency on a Tuesday afternoon. She has 1,200 completes back from a concept test for a beverage client, fielded over four days through a panel exchange, incentive of about four dollars a head. She is doing the thing she always does before the deck: reading open ends.

They are good. Not suspiciously effusive, not copy-pasted, not the word salad she learned to spot in 2019. One respondent says the packaging reminds them of a brand they drank in college and they are not sure that is a compliment. Another gives a specific, slightly grumpy note about the cap being hard to open with wet hands. A third writes two sentences about buying it for their kid's lunchbox and then hedges that the sugar content would stop them.

Every one of those is the kind of response researchers highlight in a readout. Specific. Ambivalent. Human-shaped. She flags four of them for the verbatim slide.

Here is the uncomfortable part. Nothing in that paragraph tells you whether a person wrote any of it. Every quality she just used as evidence of humanity, specificity, ambivalence, mild grumpiness, an unexpected detail about wet hands, is trivially producible by a language model that costs a fraction of a cent to run and was told to sound like a 34 year old parent in Ohio who is mildly annoyed. The signal she trusts most is the signal that became cheapest to fake.

She is not being careless. She is applying the professional judgment that worked for fifteen years. The ground moved underneath it.

Industry sources put fraudulent online survey responses somewhere between 15 percent and 40 percent, with CloudResearch citing figures in the 30 to 40 percent range and buyers' guides citing 15 to 30 percent. Every current defence is behavioural detection, which AI agents now pass. The control that does not degrade as models improve is a presence receipt bound to each response: proof that a unique human submitted it, with no identity collected and only a one way key retained.

How much survey fraud is there actually?

The honest answer is that the industry does not know, and the range of published estimates is the best evidence of that.

The incentive platform Tremendous has written that fraud cost the market research industry an estimated 350 million dollars in 2024, which it puts at roughly a tenth of total incentive spend across the industry. That is a figure about money leaving through the payout rail, which is the one thing in this system that is precisely countable, because somebody cashed a gift card.

CloudResearch, which builds participant sourcing and quality tooling and has published extensively on AI participants, has cited figures in the range of 30 to 40 percent for fraudulent online survey responses, and ran a 50,000 dollar challenge it called the Bot Olympics specifically to test detection methods against AI agents. Emporia Research, in a buyer's guide to anti-fraud tooling in sampling, puts the range at 15 to 30 percent and notes studies where advanced checks flagged close to one respondent in three as suspicious. Rep Data has discussed analysis across billions of survey starts in its fraud reporting, covered on the Greenbook podcast.

Look at what those numbers actually are. One is a payout figure. One is a vendor's estimate of a rate. One is a range from a buyer's guide. One is a flag rate from detection tooling, which is not the same thing as a fraud rate, because a flag rate includes false positives and misses whatever the tool cannot see.

Why the spread matters more than the midpoint

It is tempting to average these and say a quarter. Do not. The spread from 15 to 40 percent is not measurement noise around a true value. It is four different instruments pointed at four different things, and none of them can see the category that matters most, which is the fraud that passes.

Every one of those estimates is produced by a detector. A detector reports what it caught. The quantity a researcher actually needs is what got through, and by construction no detector can report that number. This is the same epistemic hole that makes it hard to state a true rate for any fraud category where the defence is probabilistic, and it is why you should treat every survey fraud statistic, including the ones in this post, as a lower bound on a quantity nobody has measured directly.

Why did open ended answers stop working as a quality signal?

For most of the history of online research, the open end was the researcher's most reliable instrument. It was expensive to fake at scale. A click farm could straightline a grid, could pass an attention check by reading four words, could randomise its way through a MaxDiff. What it could not cheaply do was write eighty words that sounded like a specific person having a specific mildly conflicted opinion about a beverage cap.

That was never a deep truth about human cognition. It was an economic fact about the price of fluent text. Fluent, contextually appropriate, persona consistent text used to cost roughly what a human minute cost. Now it costs less than the electricity to serve the page.

When the price of the hardest-to-fake signal collapses, the entire quality stack inverts. The checks that survive are the cheap mechanical ones, speeders, duplicate IPs, impossible device combinations, and those were always the checks that caught the least sophisticated fraud. The check that caught the sophisticated fraud is now the one that catches nothing.

The lock and the locksmith

Think of it like a lock that was secure for thirty years because picking it required a rare tool that only a few hundred people owned. The lock did not get worse. The tool got mass produced and shipped to everyone for free. No amount of admiring the lock's engineering changes what happened to your door.

The precise version: open end quality was a proxy for effort, and effort was a proxy for humanity. Language models broke the second link. Effort is no longer evidence of a person, because the effort is now nearly free and the output is indistinguishable at the level a human reviewer or a text classifier can operate on. Researchers who keep reading verbatims for authenticity are calibrating an instrument against a reference that has moved.

What do the economics of fake responses look like?

This is the part that makes the problem structural rather than a nuisance that better tooling will grind down. Work the arithmetic.

A typical consumer survey pays somewhere between one and eight dollars for eight to fifteen minutes of attention. Call it four dollars for ten minutes. For a real respondent, that is roughly 24 dollars an hour of gross incentive, which is a mediocre but real wage in some markets and a good one in others. That price was set to be worth a human's time.

Now price the same completion for an operator running language models. A ten minute survey is a few thousand tokens of input and output. At current commodity inference prices, generating a full set of coherent, persona consistent answers costs a small fraction of a cent. Add proxy costs, panel account costs amortised over many completions, and some human time to set up and maintain the operation. The marginal cost per fake completion sits well under a cent even with generous assumptions.

So the operator is buying four dollars for less than one cent. That is a margin above 99 percent, and unlike most fraud it does not require a victim to be deceived in the moment, only a filter to be passed. Supply is limited by panel account availability and detection pressure, not by labour.

PartyCost per completionRevenue per completionMargin
Honest respondent10 minutes of attentionAbout 4 dollarsA modest hourly rate
Click farm, human operated2 to 4 minutes of low wage labourAbout 4 dollarsMeaningful but bounded by labour
Agent operatedWell under one cent of inferenceAbout 4 dollarsAbove 99 percent, bounded only by account supply

Read the last row again. When a fraud has a margin above 99 percent and no labour constraint, detection does not remove the incentive, it only sets a price on entry, and an operator with that margin can pay it and keep going. Suppression works when it pushes the attacker's cost above their revenue. Here you would need to raise the cost of a fake completion several hundredfold using tools that also inconvenience real respondents. That is not a fight detection wins.

Why does every existing defence fail?

Panels are not lazy about this. The typical quality stack has five or six layers, and each of them was a reasonable answer to the fraud of its era. Go through them honestly, because understanding why each one fails is what tells you what the replacement has to look like.

Attention checks

The classic instructional manipulation check asks the respondent to select a particular option to prove they are reading. This catches inattentive humans and crude scripts. A language model reads the instruction and follows it, because following instructions embedded in text is the single thing it is best at. Attention checks now function primarily as a tax on distracted real people.

Speeders, straightliners and pattern filters

Timing and pattern filters catch operators who optimise for throughput, and cost one line of configuration to defeat: add a randomised delay from a plausible human distribution and vary grid responses with noise. The filter still removes genuine fast readers.

Device fingerprinting and proxy detection

Fingerprinting infers whether many sessions share a machine, and proxy detection flags traffic from data centres. Residential proxy networks and commodity browser automation defeat both, and they are sold as services. It is worth noting that the bot defence industry has been candid about this trajectory. hCaptcha's own materials have acknowledged that traditional fingerprints are becoming less useful as browser makers restrict them and attackers emulate them. When a detection vendor writes the obituary for its own signal, believe them.

Open end AI text detectors

This is where researchers place the most hope and where the ground is least stable. AI text classifiers have meaningful false positive rates, they degrade when a model is prompted to vary style, and they are trivially defeated by asking the model to write like a tired person on a phone. They also disproportionately flag non-native English speakers, which means the tool you deploy to remove bots quietly removes a demographic. That is a data quality problem wearing a data quality solution's clothes.

Duplicate detection and incentive holdbacks

Deduplication catches one identity reused within one panel. It cannot see across panels, and an operator with a hundred accounts is not duplicating anything. Holdbacks shift cash flow risk onto honest respondents without changing what the reviewer can see.

Notice the shape common to all six. Each one observes the response or the device and infers a probability of humanity. Each has a false positive rate that costs you real respondents. And each one's effectiveness degrades as models improve, because the thing being observed is exactly the thing models are getting better at producing.

DefenceCatchesMissesCost to the honest respondentDegrades as models improve
Attention checksInattentive humans, crude scriptsAny instruction-following modelReal people fail them and get screened outYes, already broken
Speeder and straightliner filtersThroughput-optimised operatorsAnyone who adds randomised delayFast readers removedYes
Device fingerprintingReused machines, data centre trafficResidential proxies, real browsersPrivacy-hardened users flaggedYes
AI text detection on open endsDefault-voice model outputStyle-prompted outputNon-native speakers flaggedYes, rapidly
Duplicate detectionRepeat identities in one panelMany accounts, cross-panel activityHouseholds sharing a deviceNeutral
Presence receipt per responseAny submission with no live humanA human who pastes model outputOne short interactionNo

That last row is the argument of this post, and the honest limit in its "misses" column is the sentence you should hold onto. We will come back to it and give it the space it deserves.

What is actually changing in the industry right now?

Two events from the last year tell you the market has already priced this in, even if individual buyers have not.

The first is Amazon Mechanical Turk, the substrate of a generation of behavioural science, which closed to new customers and has been reported as winding down entirely in 2026. Prolific, a competitor, framed the closure as a symptom of a data integrity problem rather than a business decision in isolation. The polite reading of an exit like that is strategic focus. The less polite reading, discussed openly in the industry, is that the unit economics of policing a low-priced open marketplace against automated participants stopped working.

The second is that Prolific shipped authenticity checks aimed specifically at distinguishing human participants from agentic AI, and published its own testing of how those checks compare against other methods. Read that as a signal about where the frontier is: a serious research platform now considers "is this participant an AI agent" a first-class product problem rather than an edge case.

Meanwhile the academic literature has been catching up. Work published by Veselovsky and colleagues in 2023, circulated on arXiv, estimated that a substantial share of crowd workers on a text summarisation task used large language models to complete it. More recently, a paper in a SAGE methods journal argued directly that chatbots are undermining crowdsourced research in the behavioural sciences. And the problem has reached peer review itself: a 2026 article in Nature reported that 21 percent of manuscript reviews submitted to a major AI conference were found to be AI generated, according to analysis by Pangram Labs. If reviewers of AI papers are outsourcing their reviews to AI, the assumption that a paid task produces human output is not a safe default anywhere.

What would actually work?

Start by stating the requirement precisely, because most of the confusion in this area comes from vague goals.

A researcher does not need to know who a respondent is. In most studies they must specifically not know, because the consent form promised anonymity and the ethics board approved it on that basis. What a researcher needs is two facts: that a live human produced this submission, and that this human has not already submitted to this study. Presence and uniqueness. Not identity. That separation, proving something true about a person without learning who they are, is the same one that makes age verification without an ID upload possible.

That distinction is why the obvious answers are wrong. A government ID gives you identity you did not want, cannot lawfully retain for most study designs, and must now defend against breach. A phone number is a weak identifier costing cents in bulk. A biometric enrolment in a central database is the worst version: a permanent record of people's faces, held by a research vendor, for the purpose of asking them about shampoo.

Presence and uniqueness without identity

Here is the construction that gives you what you need and nothing you do not.

At study entry, the respondent completes a short interaction on their own device. A face match runs entirely in the browser, on the device, and produces a vector that never leaves it. From that vector, combined with a per-study salt, the device derives a one way key. The key is stable for the same person on the same study and reveals nothing about the face it came from, in the same way a password hash reveals nothing about the password. The panel stores the key. It does not store the image, the vector, or anything that can reconstruct either.

When the respondent submits, their device signs the submission with a passkey, and the signature covers a hash of the response payload together with the study identifier and the one way key. The result is a receipt: a small signed object saying a live human, identified only by an opaque key that means nothing outside this study, submitted this specific set of answers at this time.

Uniqueness falls out for free. If the same person tries to complete the study twice, their device derives the same one way key, and the panel sees a collision without ever knowing who they are. Two hundred accounts operated by one fraudster produce either one key, if one person sits behind them, or two hundred distinct live humans, which is a completely different and much more expensive business.

The client's side of this is the part that changes commercial behaviour. Because the receipt is signed and verifies against a published key, the client can check it themselves, offline, without calling the panel and without trusting the panel's assurances. A completion rate stops being a claim in a slide and becomes a set of artifacts an auditor can verify.

What does the receipt actually contain?

Concretely, this is the object. Nothing here is identity.

{
  "v": 1,
  "study_id": "bev-concept-2026-09",
  "response_hash": "sha256:9f2c41a7e8...c3",
  "human_key": "hk_7f3a91e2c48b5d60",
  "presence": { "mode": "once_beam", "liveness": true, "ts": "2026-09-10T14:22:31Z" },
  "sig": "ed25519:MEUCIQDf...",
  "key_id": "manav-2026-03"
}

Read it field by field, because every omission is deliberate.

study_id scopes everything. The one way key is derived with it as a salt, so the same person in another study produces a different key and cannot be linked across studies by the panel, the client, or us. response_hash binds the receipt to these specific answers, so you cannot detach a valid receipt and staple it to another submission, which is the first attack you would try. human_key is the uniqueness handle: an opaque value that cannot be reversed to a face, a name, an email or a device, whose only power is to collide with itself. presence records that a live check happened, when, and that it included a challenge a static image cannot satisfy.

sig and key_id make this verifiable by a third party. Verification is a signature check against a published key, with no callback to us:

from nacl.signing import VerifyKey
import json, hashlib, base64

def verify(receipt, response_body, published_key_b64):
    # 1. the receipt must be bound to this exact response
    digest = "sha256:" + hashlib.sha256(response_body).hexdigest()
    if digest != receipt["response_hash"]:
        return False, "receipt does not match this response"

    # 2. the signature must check out against the published key
    signed = json.dumps({k: receipt[k] for k in
        ("v","study_id","response_hash","human_key","presence")},
        sort_keys=True, separators=(",",":")).encode()
    vk = VerifyKey(base64.b64decode(published_key_b64))
    vk.verify(signed, base64.b64decode(receipt["sig"].split(":")[1]))

    return True, "a live human submitted this response"

Two properties are worth naming. First, the client runs this. Not the panel. That is what turns a quality claim into a quality fact. Second, it works in ten years, on an archived dataset, with the panel out of business, because it is arithmetic against a published key rather than a query against somebody's database.

The widget that produces these receipts is a drop-in gate with declarative modes, described in the developer documentation. The closest working demonstration of the shape is the humanity gate demo, and the bot job applications demo shows the same primitive applied to bulk submission fraud, which is structurally the identical problem with a different payload.

What does this cost the honest respondent?

A fair question, and the one that decides adoption.

The interaction is a few seconds on the respondent's own phone at study entry, and for long studies it can be re-checked once rather than continuously. Compare that against the current experience of an honest respondent, which is: an attention check that insults them, a timing filter that punishes them for reading quickly, a text classifier that may reject their open end because English is their second language, and a payout held for review because a fingerprinting service does not like their privacy settings.

The existing stack is not low friction. It is friction distributed invisibly and unfairly, mostly onto people who did nothing wrong, and it does not work. Concentrating the check into one short, honest interaction that actually excludes the fraud is a better trade for the respondent, not a worse one.

There is a real equity consideration underneath this and it should be stated rather than waved at. Any check that requires a specific class of device will exclude somebody. A design that requires a recent smartphone with a working front camera will systematically under-sample the poorest respondents, and in research that is not just unfair, it is a bias that corrupts the result. The design has to work on cheap hardware, has to have a fallback path, and panels adopting it should measure completion rates by demographic before and after and publish the difference. If verification shifts your sample, you have traded one data quality problem for another.

What this cannot do

This is the section that matters most, and skipping it would make the rest dishonest.

A real human who pastes model output has still signed. If a genuine person opens the survey, reads the question, asks a chatbot to write the answer, pastes it, and submits from their own device with a live face check, the receipt is valid and correct. It proves exactly what it claims, which is that a live human submitted this. It does not claim the human composed it. This is the single largest limit and it should be stated in the first paragraph of any vendor deck that sells this, including ours.

What the receipt does to that scenario is change the economics rather than eliminate the behaviour. Fake completion volume today is bounded by account supply and inference cost, which is to say barely bounded at all. Under a presence requirement, volume is bounded by the number of real humans willing to sit through a check, which is a labour constraint again. You have moved the problem from unlimited supply back to something priced in human minutes. That is a large win and it is not a solution to authorship.

The remedy for authorship is task design, not cryptography. Time-boxed responses, questions that reference something the respondent must have looked at, probes that follow up on the respondent's own earlier answer in a way that punishes generic text, and study designs that do not reward volume. Researchers know this literature. The point is that these are the right tools for that problem, and no amount of signing addresses it.

Cross-panel uniqueness is not solved today. The one way key is scoped per study, which is a privacy property and a limitation at once. It means a person cannot be tracked between studies, and it also means one person can complete the same client's study on three different panels. Solving that without creating a linkable global identifier requires nullifier schemes, which are on our roadmap and not shipped. Say so plainly rather than implying coverage that does not exist.

Enrolment is the trust bottleneck. Everything downstream assumes the check at study entry was a real live person. A sufficiently determined operator who can defeat liveness at enrolment gets a valid key. Injection attacks against verification pipelines are a real and growing category, which we cover in the camera is no longer evidence, and the defence there is device attestation rather than better video analysis.

It changes nothing about sample representativeness. Verified humans are still whoever your panel recruits. A perfectly human sample can be a perfectly unrepresentative one, and this control does not touch that.

What to do this week

  1. Ask each panel you buy from what fraction of completions they proved were human, and listen for whether the answer describes detection or proof. The phrasing of the answer tells you more than the number.
  2. Pull the last three studies and calculate what you actually paid for suspect completes: total incentive spend multiplied by your best estimate of the fraud rate. That number is your budget for fixing it.
  3. Stop treating open end quality as an authenticity signal in your internal QA guidance. Write it down and tell the team, because this is an unlearning problem and it will not happen on its own.
  4. Check what your AI text detector does to non-native English speakers in your own data. Run the flag rate by language group. If there is a gap, you are removing a demographic, not fraud.
  5. Add a verified completion clause to your next panel contract, specifying that the vendor supplies per-response receipts the client can verify independently. Even if no vendor can meet it yet, putting it in procurement moves the market.
  6. For any study whose results will drive a decision above a threshold you care about, pilot a presence gate on one cell and compare the distribution of answers against an unverified cell. If the two differ, you have just measured your contamination directly.
  7. Publish your method. Whatever you find, the industry benefits from more disclosed numbers and fewer vendor estimates, and buyers who publish get better treatment from panels.

Frequently asked questions

How much of online survey data is fraudulent? Published estimates range from about 15 percent to 40 percent. CloudResearch has cited figures in the 30 to 40 percent range, while buyers' guides such as Emporia Research cite 15 to 30 percent. Every one of these numbers comes from a detector reporting what it caught, so treat them all as lower bounds on a quantity nobody has measured directly.

Can AI agents actually complete surveys? Yes, including passing attention checks and writing plausible open ends, which is why Prolific shipped authenticity checks aimed specifically at agentic AI participants and why CloudResearch ran a funded challenge to test detection against agents. Instruction following is the capability language models are strongest at, and most survey quality checks are instructions.

Why did Mechanical Turk shut down? Amazon closed MTurk to new customers with a reported wind-down in 2026. Competitors and commentators have framed it as a symptom of a broader data integrity problem in low-priced open marketplaces rather than an isolated business decision. Policing automated participants in a marketplace priced for cents per task is economically difficult.

Does verifying respondents mean collecting their identity? No, and it should not. The construction described here does an on-device face match, derives a one way key that cannot be reversed, and stores only that key. The panel learns that a unique live human submitted, and learns nothing about who they are. For anonymous studies that is the only lawful and ethical shape.

Will this stop a person who uses ChatGPT to write their answers? No. A live human who pastes model output produces a valid receipt, because they are a live human. The receipt proves presence and uniqueness, not authorship. It raises the cost floor from a fraction of a cent to a human's time, which is a large change in the economics, and authorship remains a task design problem.

Sources

  1. Tremendous, Protecting market research data from fraud, on an estimated 350 million dollars of incentive fraud in 2024.
  2. CloudResearch, The Bot Olympics: a 50K test of AI survey fraud detection.
  3. Greenbook, What 4 billion surveys reveal about data risk, with Rep Data.
  4. Emporia Research, A buyer's guide to anti-fraud tools in market research sampling.
  5. Prolific, Authenticity checks: identifying agentic AI participants.
  6. Prolific, MTurk's closure marks the end of an era.
  7. Nature, reporting that 21 percent of reviews at a major AI conference were AI generated, per analysis by Pangram Labs.
  8. Veselovsky and colleagues (2023), Artificial Artificial Artificial Intelligence: crowd workers widely use large language models for text production tasks, preprint on arXiv. Cited by name; see arXiv for the current version.
  9. SAGE, Chatbots are undermining crowdsourced research in the behavioral sciences.
Your best quality signal became the cheapest thing in the world to fake. Stop grading the answer and start proving the answerer.