Manav.id
Research ยท 19 min read

Twenty one percent of the reviews were written by the thing being reviewed

A fifth of the reviews submitted to a major machine learning conference appeared to be machine generated. The instinct is to buy a detector and start accusing people. That instinct is wrong, and in peer review specifically it is dangerous. The problem is not that reviewers use models. It is that no act in the system carries evidence that a person stood behind it.

An area chair opens the fourth review of a submission at two in the morning, because that is when area chairs do this work. The review is three paragraphs, well organised, grammatically immaculate. It praises the clarity of the exposition. It notes that the empirical evaluation could be more comprehensive. It suggests the authors consider additional baselines. It rates the paper a five.

Every sentence is defensible. Not one of them requires having read the paper.

The area chair has a suspicion and no way to act on it. There is no evidence of misconduct, only an absence of specificity, and plenty of tired, overloaded, entirely honest reviewers write vague reviews at two in the morning too. Accusing a colleague of fabricating a review on the basis of prose style is a serious thing to do and the area chair does not do it. The review goes into the pile and gets weighed with the others.

Multiply that moment by several thousand and you have the state of peer review.

Short answer: Nature reported in 2025 that an analysis by Pangram Labs found roughly 21 percent of reviews submitted to a major AI conference appeared to be AI generated. Detection is a poor remedy here because false accusations end careers, detectors misfire on non native English writers, and review confidentiality makes adjudication nearly impossible. The durable control is a signed review: an attestation that a named human submitted this assessment, held by the editor and never shown to the author, so blind review survives.

What did the analysis actually find?

Nature reported in 2025 that controversy had erupted after roughly 21 percent of manuscript reviews for an international AI conference were found to have been generated by artificial intelligence, based on analysis conducted by Pangram Labs, a company specialising in detecting machine generated text.

Before going further, an uncomfortable observation that this post has to make about its own headline. That 21 percent figure was produced by a detector. The rest of this post argues that detectors are unreliable. Both things are true simultaneously, and the honest reading is that the figure indicates a large problem whose precise size nobody knows. It is directional evidence, not a measurement, and anybody quoting it as a hard number, including us in the title, should say so.

The finding sits alongside a related development from the same period. Systems from Sakana AI and from Intology were reported to have produced papers that passed real peer review during 2025, at workshop level and at a conference. Take those reports at their stated scope: a small number of papers, in venues with their own standards, and in at least one case with organiser involvement in the experiment. They do not establish that machine written papers routinely pass. They establish that the boundary is porous, which is enough.

So both sides of the review relationship now contain a category that did not previously exist. Some papers were not written by the people submitting them, and some reviews were not written by the people signing them. The system that connects them assumed humans on both ends and has no mechanism to check either.

Why are reviewers reaching for the model?

Any account of this that starts from reviewer misconduct will fail, both as analysis and as a basis for reform. Understand the incentives first.

Peer review is a gift economy running at a deficit

Reviewers are not paid. Reviewing is a professional obligation sustained by reciprocity: you review because others reviewed your work, and because the field only functions if somebody does it. It is genuinely honourable, and it is also entirely uncompensated labour performed on top of a full workload.

Submission volumes at major machine learning venues have grown by large multiples over the past decade, and the reviewer pool has not grown proportionally, because the pool is drawn from the same community that is producing the submissions and each person has a fixed number of hours. The arithmetic is not subtle. When submissions grow faster than reviewers, the load per reviewer rises, and it has risen a great deal.

A conscientious review of a technical paper takes hours. A reviewer with six assignments, a deadline, a teaching load and their own submissions in flight is facing something like a full working week of unpaid labour, concentrated into the same two weeks as everyone else. When a tool appears that produces a plausible review in ninety seconds, some fraction of people will use it, and the surprising thing is that the fraction is as low as it apparently is.

This is a systemic failure, not a moral one. Treating it as reviewer misconduct misdiagnoses the cause and, more practically, alienates precisely the people whose voluntary cooperation any fix depends on. You cannot enforce your way out of a volunteer shortage.

The confidentiality problem, which is separate and real

There is a second issue that gets tangled with the first and should not be. Most publisher policies that restrict AI use in review rest on confidentiality rather than on quality. A manuscript under review is a confidential document entrusted to the reviewer, and pasting it into a third party service transmits unpublished work to a party the author never agreed to. Guidance from bodies including the Committee on Publication Ethics and policies from the major publishers converge on this point.

That argument is strong and largely independent of whether the resulting review is any good. A reviewer who uses a model locally, on a machine under their control, has not breached confidentiality. A reviewer who pastes a manuscript into a hosted chat interface has, even if their review is excellent. Keeping these two questions separate makes the policy conversation much more tractable.

Why is AI detection a particularly bad fit here?

We have argued in the piece on detection debt that classification loses to generation across many domains. Peer review is the case where the argument is strongest, and the reason is not primarily about accuracy.

The cost of a false positive is a career

In most detection settings a false positive is an inconvenience: a blocked transaction, a flagged login, a support ticket. Here, a false positive means telling a named academic that they fabricated a professional judgment. That accusation, even if quietly withdrawn, attaches to a person's reputation in a small community where reputation is the entire currency. There is no equivalent of a chargeback.

Detectors misfire on non native English writers

This is the most serious specific objection and it is well established in the literature. Research on GPT detectors, including work by Liang and colleagues published in Patterns in 2023, found that detectors classified writing by non native English speakers as machine generated at substantially elevated rates, because the features detectors key on, such as lower lexical variety and more conventional phrasing, are also features of competent second language academic writing.

Consider who that lands on. International research communities. Scholars from institutions without editing support. Early career researchers writing in their second or third language. Deploying a detector across a review pool means systematically directing suspicion at those groups, and doing so under the banner of research integrity would be a genuinely shameful outcome.

Confidentiality makes adjudication nearly impossible

Suppose a detector flags a review. What happens next? In a double blind venue the reviewer is anonymous to the author, the review is confidential, and there is no forum in which the accused can defend themselves publicly without unmasking. The venue must adjudicate privately, on the basis of a probabilistic score, against a person who cannot mount a public defence, with no way to compel evidence. That is not a process anyone should want to run.

Put the three together and the conclusion is not merely that detection is ineffective. It is that detection is ethically hazardous in this setting in a way it is not elsewhere. That is a stronger claim than the usual efficacy argument and it is the one editors should weigh.

What is peer review actually built on?

Step back and ask what the system rests on, because the answer determines what a fix has to preserve.

Peer review is not an enforcement regime. There is no auditor, no penalty schedule, no compliance function. It is a trust arrangement in which qualified people agree to do unpaid work carefully, and the whole edifice of scientific credibility, tenure decisions, grant allocation and, downstream, clinical and policy guidance sits on top of that arrangement.

A system built on trust degrades differently from one built on enforcement. It does not fail loudly when a rule is broken. It fails quietly as participants observe that others are not carrying their share, and adjust. The dangerous thing about a 21 percent figure, whatever its precision, is not the reviews themselves. It is what an honest reviewer does after reading it.

So the design requirement is unusual: whatever gets added must strengthen accountability without converting a gift economy into a surveillance regime, because the second would destroy the thing it was meant to protect. That constraint rules out most of what is currently being proposed.

What would a signed review prove, and what would it not?

The proposal is narrow. When a reviewer submits, the platform additionally collects an attestation produced on a device the reviewer enrolled once: a signature over a hash of the review text, bound to the reviewer's account, produced at submission time.

{
  "venue":         "ICLR-2027",
  "submission_id": "4821",
  "review_sha256": "b71f...9ac2",
  "reviewer_ref":  "orcid:0000-0002-1825-0097",
  "submitted_at":  "2026-11-14T02:41:19Z",
  "assistance":    "local-model-drafting",
  "attests":       "read-and-endorsed",
  "sig":           "ed25519:0a4e...71bd",
  "key_id":        "manav-2026-09"
}

Read the fields honestly, because the interesting ones are the two that make no technical claim at all.

attests is a declaration by the reviewer, not a measurement. The signature does not prove the reviewer read the paper. Nothing can prove that, and any vendor claiming otherwise is lying to you. What the signature does is convert an implicit assumption into an explicit, attributable, durable statement. The reviewer is asserting something specific, with their key, and that assertion persists.

assistance is the same: a disclosure the reviewer makes, in a structured field rather than buried in a comment box. It exists because prohibition does not work and honest disclosure might. A venue that permits local model drafting but forbids uploading manuscripts to hosted services can express that policy, and reviewers can comply with it visibly.

What the signature genuinely establishes is narrow and still valuable: a specific enrolled human, at a specific time, put their name to this exact text. Not a login session that could have been shared, not an email address, not an account somebody else had access to. If that review is later found to be fabricated, the attribution is not in dispute, which changes the calculus for anyone considering it.

How do you sign a review without destroying blind review?

This is the objection that kills naive versions of the proposal, and it deserves the most careful paragraph in the post.

Double blind review exists for good reasons. Authors should not know who reviewed them, because knowing enables retaliation, ingratiation and a great deal of subtle social pressure, and it disadvantages junior reviewers reviewing senior authors. Any scheme that attached a verifiable name to a review visible to authors would demolish this, and would be correctly rejected by every venue.

The resolution is that the signature has an audience of one. The reviewing platform holds the attestation. The editor or area chair can verify it. The author receives exactly what they receive today: an anonymous review. The signature is not published, not attached to the review text shown to authors, and not disclosed at acceptance. It is an internal integrity artifact, in the same way that the platform already knows the reviewer's identity while the author does not.

This is not a compromise or a partial answer, it is simply the correct architecture. Blind review has never meant that nobody knows who the reviewer is. It means the author does not. Adding cryptographic weight to something the platform already knows changes nothing about the author's view and everything about whether the platform's knowledge is trustworthy.

There is a second, quieter benefit. Reviewing is unpaid and largely invisible work, and reviewers get little credit for it. A reviewer who accumulates verified attestations across venues holds a portable record of service that they own, which they can present to a hiring committee or a funder without the venue having to vouch for them individually. We described the general shape of this in the verified work passport piece. Give people credit for the work and some of the load problem eases at the margin, which is worth more than any enforcement mechanism.

Comparing the available measures

MeasureWhat it establishesReviewer burdenCost to anonymityRisk of unjust harm
Do nothingNothingNoneNoneNone directly, corrosion over time
Ban AI use in reviewA rule existsNoneNoneLow, but drives use underground
Honour system disclosureWhat honest people declareTrivialNoneNone
AI text detection on reviewsA probabilityNoneNoneHigh. Career damage, biased against non native writers
Open review with named reviewersFull attributionModerateTotal. Destroys blind reviewHigh, enables retaliation
ORCID linked reviewer accountsA persistent identifierLowNoneNone, but proves an account not a presence
Signed review, editor heldA named human submitted this textLow, one enrolmentNone. Author sees nothingLow, no accusation is generated

Note the column that usually goes unexamined. Every measure that generates an accusation carries a risk of unjust harm, and the signed review generates no accusation at all. It produces a record, and records do not have false positive rates.

What this cannot do

The limits are substantial and stating them is more useful than the proposal itself.

It does not detect AI assistance. Not partially, not indirectly, not at all. A reviewer who generates a review with a model and signs it has produced a valid signature over machine written text. If you want to know whether a model wrote the words, this does not tell you, and this post argues you should stop trying to find out that way.

It does not prove the reviewer read the paper. There is no technology that proves comprehension and there never will be. The attestation makes a claim attributable. It does not make it true.

Many uses of AI in review are legitimate. A non native English speaker using a model to polish their phrasing is doing something entirely proper and is arguably being disadvantaged by the current discourse. A reviewer using a model to check whether they have missed relevant prior work is doing something useful. Policy should distinguish assistance from substitution, and that distinction lives in venue rules, not in cryptography.

It adds friction to a volunteer population that is already overloaded. Asking exhausted reviewers to do one more thing is a real cost and the enrolment must be a one time action measured in seconds, or it will not be adopted, and rightly so.

Device access is not universal. Reviewers work in every country and every institutional circumstance, and any requirement must have an alternative path administered by the venue.

It does nothing about the underlying shortage. This is the most important limit. The cause of the problem is that reviewing load has outgrown the volunteer pool. A signature does not review a paper. Venues that adopt integrity measures without addressing submission volume, reviewer recognition, or compensation are treating a symptom, and the symptom will return.

What to do this week

For programme chairs, editors, and publishers.

  1. Separate the confidentiality rule from the quality rule. State plainly that manuscripts must not be uploaded to third party services, which is a clear, enforceable, well grounded rule. Then treat the question of model assisted drafting as a separate policy conversation.
  2. Do not deploy a text detector against your reviewers. If you take one thing from this piece, take that. The false positive falls on non native English speakers and the accusation cannot be fairly adjudicated under confidentiality.
  3. Add a structured assistance disclosure field. A dropdown with honest options, no penalty attached for the permitted ones. You will learn more from voluntary disclosure than from any classifier.
  4. Publish your own composition. Report how many reviews were submitted, how many reviewers were assigned, and what your load per reviewer actually was. The load number is the root cause and it is rarely published.
  5. Give reviewers something they can keep. A verifiable record of service that a reviewer owns and can show a funder or a committee costs the venue almost nothing and addresses the incentive problem directly.
  6. Pilot signed reviews on one track. Keep it editor held, do not show it to authors, and measure adoption and complaints rather than guessing.
  7. Fix the load, or accept the outcome. Desk rejection thresholds, submission caps, reviewer compensation and paid editorial support are the levers that actually move this. Integrity tooling buys time; it does not buy reviewers.

The enrolment flow is visible in the walk up demo, and the receipt format and verification are documented in the developer documentation.

Frequently asked questions

How many peer reviews are AI generated? Nature reported in 2025 that an analysis by Pangram Labs found roughly 21 percent of reviews for a major AI conference appeared machine generated. Treat that as directional rather than precise, since the figure was itself produced by a detector, and detectors have known reliability problems. No trustworthy general estimate across disciplines exists.

Can AI written papers pass peer review? Reported cases from 2025 indicate that systems from Sakana AI and Intology produced papers that passed review at workshop and conference level. Those were small numbers in specific venues, in at least one case with organiser involvement, so they show the boundary is porous rather than that machine written papers routinely succeed.

How can journals verify that reviewers are human? By collecting an attestation at submission produced on a device the reviewer enrolled, signing a hash of the review text. The editor can verify it and the author never sees it, so blind review is preserved. This proves a named human submitted the text. It does not prove they wrote it or read the paper.

Are AI detectors accurate enough to use on reviews? No, and the specific reason matters. Research including work by Liang and colleagues in Patterns found detectors flag non native English writing as machine generated at substantially elevated rates. Deploying one across a review pool directs suspicion at international scholars, and confidentiality makes any resulting accusation nearly impossible to adjudicate fairly.

Would signing reviews break double blind review? No, provided the attestation is held by the venue rather than shown to authors. The platform already knows who the reviewer is. A signature adds cryptographic weight to knowledge the platform has and changes nothing about what the author sees.

Is it wrong for a reviewer to use AI at all? Not necessarily, and the debate conflates two questions. Uploading a confidential manuscript to a hosted service breaches confidentiality regardless of the review's quality. Using a local model to polish phrasing, particularly for a second language writer, is a different act. Venue policy should address them separately.

Does ORCID already solve this? Partly. ORCID provides a persistent researcher identifier, which is genuinely useful for linking a person to their record. It establishes that an account exists and is associated with a researcher. It does not establish that the person was present and acting at the moment a specific review was submitted.

Sources

  1. Nature news, 2025, reporting that 21 percent of manuscript reviews for an international AI conference were found to be AI generated, based on analysis by Pangram Labs: nature.com
  2. Liang et al., GPT detectors are biased against non native English writers, Patterns, 2023: cell.com/patterns
  3. Committee on Publication Ethics, guidance on artificial intelligence and peer review: publicationethics.org
  4. International Committee of Medical Journal Editors, recommendations covering AI use in submission and review: icmje.org
  5. Springer Nature, editorial policies on AI use by reviewers and authors: springernature.com policies
  6. Elsevier, publishing ethics policies including reviewer use of generative AI: elsevier.com policies and standards
  7. OpenReview, the reviewing platform used by several major machine learning venues: openreview.net
  8. ORCID, persistent identifiers for researchers and their contributions: orcid.org
No technology proves that a reviewer read the paper. What a signature does is make the claim that they did into something a person actually made, attributably, rather than something the system quietly assumed.