The human feedback training your model may not be human
Reinforcement learning from human feedback rests on one assumption nobody verifies: that the human was human. Published research has found crowd workers routing text tasks through language models, a major conference found a fifth of its reviews machine generated, and the platform that created this market is shutting down over data integrity. Here is what a per judgment receipt would change, and the one thing it cannot fix.
It is 11:40pm and the queue says forty one tasks remaining. The pay is good, better than good for this kind of work, and the batch closes at eight in the morning.
The task is to read two model answers to a question about renal dosing in an elderly patient, decide which one a physician would prefer, and write two sentences saying why. The annotator is a nurse practitioner. That is why they were recruited onto this project, and it is why they are paid roughly four times the rate of a general labelling queue.
They also have a second browser window open. In it is a chatbot that is quite good at comparing two answers about renal dosing and writing two sentences about why one is better.
They are not the villain of this story. On task one they read the chatbot's reasoning closely, disagree with part of it, and edit. On task nineteen they read it, agree, and submit. By task thirty they have stopped reading it. The work is still accurate, or accurate enough, and it is certainly more consistent than it was at 11:40pm.
Those judgments enter a preference dataset the following week. They are indistinguishable from the ones before them. In three months they will help define what a frontier model believes a physician prefers.
Is reinforcement learning from human feedback actually produced by humans? Increasingly, nobody can prove it. Published research has found crowd workers routing text tasks through language models, Amazon Mechanical Turk is closing amid data integrity concerns, and AI text detectors fail on paraphrase. Agreement metrics cannot distinguish two humans from two models. The only durable answer is a receipt bound to each submitted judgment, proving a present human produced it.
What does a contaminated judgment actually look like?
The intuition most people have is that machine generated annotation looks worse than human annotation. Sloppier, more generic, easier to spot. That intuition is backwards, and the fact that it is backwards is the whole problem.
A model routed judgment is usually more fluent than the median human judgment. It is better spelled. Its two sentence rationale is better structured. It arrives faster and it is more internally consistent across a batch. On every proxy a data platform actually measures, it looks like excellent work.
There are three distinct routes into a dataset, and they need to be separated because they call for different responses.
The honest worker with a second window
This is the nurse practitioner above, and it is by far the most common case. A real, qualified, recruited person is present, doing the task, and using a model as an accelerant. Early in a batch they are genuinely reviewing. Late in a batch, under deadline, review decays into acceptance. There is no moment where they decide to commit fraud. There is a gradient, and the platform has no instrument pointed at it.
The rented account
A worker who passed identity checks, built a reputation score, and qualified for premium projects sells or rents access to that account. The buyer may be in a different country, may have none of the credentials the project required, and may be running several accounts at once. Every reputation signal the platform trusts is now pointing at the wrong person. Reputation tiers, which platforms lean on heavily, are precisely the asset being resold.
The automated account
The fully automated case, where a script or an agent completes tasks with no person present, is the rarest and the easiest to catch, because it tends to be fast and regular in ways that look mechanical. It is also the case that everyone builds defences against, which is why it accounts for the least contamination. Defences get built against the failure mode that is easy to imagine.
The uncomfortable summary: the industry has invested in stopping route three, has partial coverage of route two, and has almost nothing pointed at route one, which is where most of the contamination comes from and where the workers are the most expensive.
How much of the human feedback is actually human?
Nobody knows precisely, which is itself the finding. What exists is a set of independent measurements from different corners of the same problem, and they point the same way.
The most cited direct measurement is a 2023 study from EPFL, published as Artificial Artificial Artificial Intelligence: Crowd Workers Widely Use Large Language Models for Text Production Tasks by Veselovsky, Ribeiro and West. Using a combination of keystroke analysis and synthetic text classification on a text summarisation task run on Mechanical Turk, the authors estimated that between roughly a third and a half of the crowd workers in their sample used large language models to complete the work. That study is now three years old, it covers one task type on one platform, and the tools available to workers have improved considerably since. Treat it as a floor, not a ceiling, and treat its precision with care, since estimating this at all requires the same detection methods the rest of this post argues are unreliable.
The second measurement comes from scholarly peer review, which is a useful proxy because reviewers are far more qualified and far more accountable than crowd workers. In 2026 Nature reported that 21 percent of manuscript reviews submitted to a major machine learning conference were assessed as AI generated, based on analysis by Pangram Labs. These are named academics attaching their professional reputation to the output. If a fifth of that population routes the work through a model, the base rate for anonymous piecework paid by the task is not going to be lower.
The third is methodological. Researchers who depend on crowdsourced behavioural data have been publishing warnings for two years. Work published in SAGE journals has examined how chatbots undermine crowdsourced research in the behavioural sciences, describing exactly the pattern above: data that passes attention checks, reads well, and no longer measures what the instrument was designed to measure.
The fourth is commercial, and it is the loudest signal of all. Amazon Mechanical Turk, the platform that effectively created this market, closed to new customers and announced its shutdown. Prolific, a competitor, framed the closure as a symptom of a broader data integrity problem rather than a business decision, which is a striking thing for a competitor to say, since the generous reading costs them the easy narrative. Prolific has separately shipped authenticity checks aimed specifically at distinguishing human participants from agentic AI, which tells you what they believe the threat model now is.
The fifth is the survey research industry's estimate of itself. Depending on methodology and how strictly fraud is defined, industry sources put the fraudulent or unusable share of online panel responses somewhere between roughly 15 and 40 percent, and CloudResearch has run public challenges testing whether current defences catch AI agents completing surveys. The width of that range is the point: an industry that cannot state its own contamination rate to within twenty five percentage points does not have an instrument.
Five independent vantage points. None of them is a clean number. All of them say the same thing: the assumption that a human label came from a human is no longer funded by evidence.
Why is preference contamination worse than model collapse?
These get conflated constantly, and separating them is the most useful thing this post can do for someone who builds models.
Model collapse is a problem with the distribution
Model collapse, in the sense that has been studied and named over the last few years, is what happens when models train on the outputs of previous models. The tails of the distribution thin out. Rare constructions disappear. Diversity narrows generation after generation. It is a real problem, it is measurable, and importantly it is a problem in the training data, which is the part of the pipeline everyone is already watching.
Preference contamination is a problem with the objective
Human feedback does something different from training data. It does not teach the model what text looks like. It teaches the model what better means. Preference pairs train a reward model, and the reward model is the thing that defines the target the policy optimises toward. It is not one input among many. It is the specification.
So when preference data comes from a model, you are not adding a little synthetic text to a large corpus. You are letting a previous model write the definition of the thing you are trying to improve on. The next model then optimises hard against that definition, and it will succeed, and what it succeeds at is matching a previous model's approximation of human preference.
Here is the analogy worth keeping. Until 2019 the kilogram was defined by a physical object in a vault outside Paris, and every calibrated scale in the world traced its accuracy back through a chain of comparisons to that one lump of metal. Now imagine that chain quietly breaks. Scale A is calibrated against scale B, which was calibrated against scale C, and none of them has touched the reference in years. Every scale still agrees with every other scale. They agree beautifully. Their readings are precise, repeatable, mutually consistent, and collectively drifting away from what a kilogram is, and there is nothing in the readings that could tell you.
Human preference is the reference kilogram of alignment work. It is the only thing in the loop that is not derived from something else. If the chain back to it breaks, every downstream metric continues to look excellent.
Why the error is silent, and why it looks like an improvement
This is the part that should genuinely worry a research lead.
The main quality instrument on annotation pipelines is inter annotator agreement. Two or more people label the same item and you measure how often they concur. High agreement is treated as evidence of a clear rubric and careful work. Low agreement triggers review.
Now consider what happens as model routing spreads through a worker pool. Two humans reading a genuinely ambiguous clinical preference item will disagree a meaningful fraction of the time, because the item is ambiguous and humans differ. Two humans who both routed the item through similar models will agree far more often, because they are effectively sampling the same distribution twice.
Contamination raises inter annotator agreement. The primary quality metric moves in the direction that reads as improvement. A pipeline getting quietly worse produces a dashboard that is getting better, and the better it looks the less scrutiny it attracts. There is no alarm to miss, because the alarm is wired backwards.
This is the same structural failure that shows up across the identity problems collected in the Identity Failure Map: the system measures a proxy, the proxy is exactly what the failure mode optimises, and the measurement gets more confident as the ground truth degrades.
Why do the current defences fail?
Every defence in production today is a form of detection, and detection has a specific structural weakness here: the thing you are trying to detect is produced by systems that improve faster than detectors do, and is then filtered through a human who can smooth over whatever tells remain.
| Defence | What it actually proves | How it fails |
|---|---|---|
| Gold standard questions | The submitter can answer items with known answers | Models answer gold items well, frequently better than tired humans. Seeding more gold items screens out humans faster than machines. |
| Inter annotator agreement | Two submitters concurred | Two models concur more than two humans. Contamination raises the metric, so the dashboard improves as quality falls. |
| AI text classifiers | Text resembles a distribution the classifier was trained on | Light paraphrase defeats them. False positives fall hardest on non native English speakers, so enforcement is both unreliable and unfair. |
| Keystroke and timing analysis | Input arrived with human like rhythm | Retyping defeats it entirely. It is continuous monitoring, which drives away the credentialed experts the premium projects exist to recruit. |
| Reputation tiers | An account has a long clean history | Account history is the exact asset rented and resold. Reputation makes a compromised account more valuable, not less. |
| Video proctoring | A face was visible during the session | Expensive per hour, invasive, and camera injection attacks against verification pipelines are now well documented by identity vendors. |
| Per judgment human receipt | A present, unique human submitted this specific judgment | Does not prove the human wrote the text unaided. Requires enrolment and a device. |
Look down the second column. Six of the seven rows prove something adjacent to the question. Only the last row proves a fact about the submission itself, and it is the only row whose failure mode is a limitation rather than a defeat.
What would proof of human work actually look like?
The shape of the fix follows from the shape of the failure. If the missing fact is "a specific present human produced this specific judgment", then the artifact you need is a proof attached to the judgment, created at the moment of submission, verifiable later by someone who was not there.
Concretely: the worker enrols once on a device they control. Enrolment binds a key to a present human using an on device face match with a liveness challenge, where the face never leaves the device and only a one way key is retained. There is no biometric template in a vault anywhere, which matters both ethically and because a vault is a liability that eventually leaks.
From then on, submitting a judgment produces a signature over that judgment.
What actually gets signed
The signature covers a canonical payload, not a vague session claim. Concretely:
{
"project": "clinical-preference-v3",
"task_id": "pref-2026-09-08-41837",
"submission_digest": "sha256:9f2c41ab...e77d",
"submitted_at": "2026-09-08T23:41:07Z",
"worker_key": "z6MkfR8hV2...pQ",
"presence": "beam-attested",
"presence_age_s": 412
}
Three fields carry the weight. The submission_digest binds the receipt to this exact set of bytes, so a receipt cannot be lifted from one judgment and pasted onto another. The worker_key is a pseudonymous public key, not a name, not an email, not a government identifier, so the platform can prove uniqueness and presence without learning who the worker is. The presence_age_s records how long ago the presence check was satisfied, so a buyer can set their own freshness policy rather than inheriting the platform's.
How a buyer checks it
The verification is deliberately boring, and boring is the feature. It runs offline, against a published key, with no call back to the party that issued it.
from manav import verify_receipt, canonical
def audit(dataset):
total = fresh = 0
for row in dataset:
total += 1
r = row.get("human_work_receipt")
if not r:
continue
# 1. does the receipt actually cover THIS submission?
if r["submission_digest"] != canonical.digest(row["submission"]):
continue
# 2. is the signature valid against the published key?
if not verify_receipt(r): # offline, no network call
continue
# 3. does presence meet OUR freshness policy, not the vendor's?
if r["presence_age_s"] <= 900:
fresh += 1
return fresh / total # human verified share
The number that falls out of the bottom is the one that does not currently exist anywhere in this industry: the human verified share of a dataset. Today a lab buying preference data receives a spreadsheet and a contractual assurance. With receipts it receives a number it can compute itself, from the data, without trusting the vendor's word or the vendor's monitoring.
Two properties are worth dwelling on because they are what make this acceptable to the workers as well as the buyers. The platform proves a human was present without knowing which human, because the key is pseudonymous. And it proves presence without observing the work, because nothing about keystrokes, screen contents, timing patterns or webcam feed is captured or transmitted. This is the opposite of the monitoring approach: it produces more certainty for the buyer and less exposure for the worker.
The receipts also belong to the worker. Accumulated in a wallet they control, a portable record of verified human work is exactly the asset described in the verified work passport, and it is the mechanism underneath proof of human work. A worker who has built a receipt history at one vendor can present it at the next one instead of starting from zero, which is the first thing in this market that has ever given a good annotator portable leverage.
What this cannot prove
Here is the sentence that matters most in this post, and it needs to be stated plainly rather than buried in a limitations paragraph at the end.
A human who pastes model output has still signed it. A per judgment receipt proves human presence and human accountability. It does not prove human origination of the text. The nurse practitioner in the opening scene would produce a perfectly valid receipt on task thirty, when they had stopped reading.
Anyone who tells you cryptography closes that gap is selling something. Closing it is a task design problem, and the honest set of tools is different: structure tasks so that a model's answer is not directly usable, require judgments that reference specifics only a present reader could supply, use response time distributions as a design signal rather than an enforcement one, and pay well enough that the deadline pressure driving route one contamination does not exist.
What the receipt does change is the base of the pyramid. It removes the rented account, because the enrolled human must be present at submission. It removes the automated account entirely. And it converts route one from an invisible gradient into an accountable act by a known, present, pseudonymous party who can be excluded from a project if their work fails quality review. You cannot exclude a farm you cannot see.
Three further limits worth naming:
- Device requirements. Enrolment needs a device with a modern authenticator, which is a real inclusion question for a global annotation workforce and will exclude some willing workers on older hardware in lower income markets. That cost is genuine and should be weighed, not waved away.
- Long tasks. Expert review sessions can run for hours. A presence check at submission proves presence at submission. Attesting continuously would be monitoring, which defeats the purpose, so the honest design is periodic re attestation with the cadence disclosed to the buyer, and the
presence_age_sfield exists precisely so the buyer can judge for themselves. - Quality is a separate question. A receipt says a human was there. It says nothing about whether that human was any good. Gold items and agreement metrics still have a job, they are just no longer being asked to do a job they were never able to do.
There is also a case the industry keeps treating as fraud that is not. Model assisted labelling is often legitimate and sometimes explicitly desirable, particularly for first pass annotation that a human then corrects. The problem is not that a model touched the work. The problem is that nobody can tell the difference between a model assisted judgment and a human judgment, so both get priced and trusted identically. The honest fix is to record model assistance as what it is, delegated work, using the same delegation chain machinery described in the authority graph, and let buyers price the two differently. That is a better outcome than a ban nobody can enforce.
What would a Human Feedback Integrity Benchmark measure?
The measurement that would move this industry does not exist yet, and it is buildable. Call it the human verified share: for a given dataset, the fraction of judgments carrying a valid, fresh, submission bound human receipt.
It is a good benchmark for the same reasons the useful benchmarks in security are good. It is computed from the artifact rather than asserted by the producer. It is verifiable offline by anyone holding the dataset. It degrades gracefully, since a dataset at 60 percent is meaningfully better than one at 0 percent and both are publishable. And it creates a price signal, because once two vendors quote the same rate and one can show 95 percent while the other cannot show anything, the market does the rest.
The regulatory pull is arriving on its own schedule. The EU AI Act's data governance obligations in Article 10 require providers of high risk systems to apply governance practices covering the origin and suitability of training, validation and testing data. "We bought it from a reputable vendor" is going to age badly as an answer to a question about provenance. A computable share is an answer that survives an auditor.
What to do this week
- Ask your vendor one question and write down the answer: "for what fraction of the judgments in our last delivery can you prove a human was present at submission?" The response, including any discomfort, is the most informative thing you will learn this quarter.
- Plot inter annotator agreement over the last eighteen months. If it has been drifting up on subjective or ambiguous items, do not celebrate. Treat rising agreement on genuinely hard items as a contamination indicator until you can rule it out.
- Re label a small gold set with staff. Take 200 items from a recent premium delivery, have your own qualified people label them independently, and compare distributions rather than accuracy. Contamination shows up as reduced variance more clearly than as error.
- Separate your spend by contamination cost. Not all feedback matters equally. Preference data feeding a reward model deserves provenance requirements that bulk classification does not. Rank your projects by what a bad judgment costs downstream.
- Stop paying for AI text detection on annotator output. It does not work on paraphrase, it misfires on non native speakers, and the budget is better spent on task design or on presence proof.
- Put a provenance clause in the next contract. Ask for per judgment receipts, or at minimum for the vendor to state their measured human verified share and the method behind it. Contracts move markets faster than blog posts.
- Audit the task design on your highest value project. For each task type ask a blunt question: could a competent model complete this from the prompt alone? If yes, no verification scheme will save the data, and the redesign is the actual fix.
- Look at what a receipt looks like in practice in the presence tracker demo, and at how the signature and verification work in the developer documentation.
The correction that is coming either way
There is a version of the next two years where this resolves badly. Labs conclude that human feedback is unverifiable, decide synthetic preference data is close enough, and quietly stop paying for the real thing. The reference kilogram gets thrown away because nobody could prove which scales still touched it. Rates collapse, the good annotators leave, and alignment research spends a decade optimising against its own reflection.
There is another version. Provenance becomes a purchased property rather than an assumed one, datasets ship with a computable human verified share the way software ships with a bill of materials, and vendors who can prove presence charge more and pay more. Farms get priced out rather than policed out. The nurse practitioner at 11:40pm is paid enough, on a task designed well enough, that the second window stays closed.
The difference between those futures is not better detection. It is whether the industry decides that "human feedback" should be a claim someone is required to back.
Frequently asked questions
Is RLHF data actually produced by humans? Increasingly, nobody can prove it either way. A 2023 EPFL study estimated that between roughly a third and a half of crowd workers on a text summarisation task used language models to complete it, and Nature reported in 2026 that 21 percent of reviews at a major machine learning conference were assessed as AI generated. Neither agreement metrics nor text classifiers can settle the question, because both measure proxies rather than the submission itself.
Why is Amazon Mechanical Turk shutting down? Amazon closed the platform to new customers and announced a shutdown. Competitors have framed the closure as a symptom of a broader data integrity problem in crowdsourced work rather than purely a business decision. Whatever the internal reasoning, the closure of the platform that created this market is a strong signal about how hard verifying human participation has become at scale.
Can AI text detectors catch annotators who route work through a model? Not reliably. Light paraphrase defeats current classifiers, and false positives fall hardest on non native English speakers, which makes enforcement both inaccurate and unfair. Detection also faces a structural problem: the generating models improve faster than the detectors, and a human sitting in the loop can smooth over whatever statistical tells remain.
Does a human presence receipt prove the annotator did not use AI? No, and this is the most important limitation to understand. A receipt proves a specific present human submitted a specific judgment. A human who pastes model output has still signed it. The receipt removes rented and automated accounts entirely and makes the remaining behaviour accountable to a known party, but closing the paste gap is a task design problem, not a cryptography problem.
Is requiring proof of human work a form of surveillance? It is designed to be the opposite. Nothing about keystrokes, screen contents, timing patterns or webcam feed is captured or transmitted. Enrolment uses an on device face match where the face never leaves the device and only a one way key is retained, and the key that signs each judgment is pseudonymous, so a buyer can prove a human was present without learning who that human is.
What is the human verified share of a dataset? It is the fraction of judgments in a dataset that carry a valid, fresh receipt bound to that exact submission. It can be computed by the buyer, offline, from the data itself, without trusting the vendor's assurances. No equivalent measure exists in the market today, which is why a lab buying preference data currently has to accept a contractual promise instead of a number.
How does this relate to the EU AI Act? Article 10 of the EU AI Act requires providers of high risk systems to apply data governance practices covering the origin and suitability of training, validation and testing data. A computable provenance figure is a far stronger answer to a regulator's question about data origin than a statement that the data was purchased from a reputable vendor.
Sources
- Veselovsky, Ribeiro and West, Artificial Artificial Artificial Intelligence: Crowd Workers Widely Use Large Language Models for Text Production Tasks, EPFL, 2023. arxiv.org/abs/2306.07899
- Nature news, on AI generated peer reviews at a major machine learning conference, 2026, citing analysis by Pangram Labs. nature.com
- Advances in Methods and Practices in Psychological Science, on chatbots and crowdsourced behavioural research. journals.sagepub.com
- Prolific, on the closure of Amazon Mechanical Turk and the underlying data quality challenge. prolific.com
- CloudResearch, public testing of survey fraud defences against AI agents. cloudresearch.com
- EU AI Act, Article 10, data and data governance. artificialintelligenceact.eu
Human preference is the reference kilogram of alignment work. If the chain back to it breaks, every downstream metric keeps looking excellent.