Manav.id
Regulated ยท 19 min read

Twenty two million comments. Eighteen million of them were fake.

The most consequential thing about the FCC net neutrality docket was not that software wrote comments. It was that millions of them were filed under the names of real people who never wrote a word. That is impersonation, and the fix for impersonation is not asking everyone else to prove who they are.

Imagine opening a search engine, typing your own name out of idle curiosity, and finding a strongly worded opinion about telecommunications policy submitted to a federal agency, attributed to you, at your home address, on a Tuesday you spent doing something else entirely.

You did not write it. You have no particular view on Title II reclassification. You are not sure what Title II is. But there it sits in a public federal docket, in your name, and it will sit there permanently, because agency comment records are not deleted.

This happened to a great many people. The New York Attorney General's 2021 investigation into the Federal Communications Commission's net neutrality proceeding found that millions of comments had been submitted using the identities of real Americans without their knowledge or consent, drawn from data held by lead generation companies. The investigation also documented submissions made in the names of deceased people, which is a detail that tends to stay with you.

The story is usually told as a bot story. It is more accurately an identity theft story that happened to be conducted at scale through a comment portal, and telling it correctly matters, because the two versions point at completely different remedies.

Short answer: Of the roughly 22 million comments filed in the FCC's 2017 net neutrality proceeding, the New York Attorney General found that nearly 18 million were fake, with millions submitted under real people's identities without consent. Generative text has since defeated the duplicate detection agencies relied on. The fix that survives constitutional scrutiny is not identity collection but a one human, one comment presence proof that reveals no name, combined with clear disclosure of which submissions carry one.

What actually happened in the FCC net neutrality docket?

In 2017 the FCC opened a proceeding to reconsider its approach to broadband regulation. Public interest was enormous, and the docket received more than 22 million comments through the Commission's Electronic Comment Filing System, which was and remains a public system that anybody can file into.

The numbers

The New York Attorney General opened an investigation and published its findings in a 2021 report on fake comments in the proceeding. The central finding was that nearly 18 million of the comments were fabricated. The report described a broadband industry funded campaign that generated millions of comments through lead generation firms, and, separately, a campaign on the other side of the policy question, reportedly the work of one college student, that generated millions more in support of net neutrality using invented identities.

Read that again, because the symmetry is the part people forget. This was not one side cheating. Both directions of the policy argument were represented by industrial quantities of comments that nobody had written. The Attorney General reached agreements with several of the lead generation companies involved, with financial terms that have been reported variously and which we will not restate precisely here.

The part almost everyone gets wrong

The headline number, 18 million fake comments, gets repeated as though the problem were volume. It was not. Volume is a nuisance. The legally serious part, the part that produced an Attorney General investigation rather than a shrug, was that a large share of those comments were falsely attributed. Real names, real addresses, real people, harvested from marketing databases and attached to political opinions those people had never expressed.

That distinction determines everything downstream. If the problem is that machines can write text, then the remedy sounds like detecting machine written text, which is a losing arms race we have described at length in the piece on detection debt. If the problem is that anybody can file a statement in your name, then the remedy is a way for a filing to carry evidence that the person named actually made it. Those are not the same project, and the second one is tractable.

The Administrative Conference of the United States, which advises federal agencies on administrative procedure, took up exactly this in its Recommendation 2021-1, addressing mass, computer generated, and falsely attributed comments. The fact that a body of administrative law experts felt the need to name falsely attributed comments as a distinct category, separate from mass comments and separate from computer generated ones, tells you the taxonomy matters.

Why does a fake comment matter if rulemaking is not a vote?

This is the strongest objection to everything that follows, and it deserves a serious answer rather than a dismissal, because it is substantially correct.

Notice and comment rulemaking under the Administrative Procedure Act is not a referendum. When an agency proposes a rule, it must give notice, accept comment, and then consider the relevant matter presented before issuing a final rule. Courts reviewing that process ask whether the agency engaged with significant comments and gave a reasoned explanation. They do not ask which side submitted more postcards. An agency that changed its position because it received nine million comments rather than four million would be doing something closer to arbitrary than to reasoned.

So in the narrowest legal sense, a flood of fabricated comments saying nothing substantive should not change a rule. Fine. Here are the three harms that persist anyway.

Harm one: the record is the record

An agency must build an administrative record and defend it in court. When a substantial fraction of that record is fabricated, every actor downstream inherits a contaminated artifact. Agency staff must process it. Litigants cite it. Journalists characterise it. Researchers study it. And a court reviewing the agency's reasoning is looking at a record whose composition is unknown. Nobody in that chain has a way to separate the genuine from the manufactured, which means the genuine comments are devalued along with the fake ones.

Harm two: it costs real money and real attention

Comment processing is not free. Agencies deduplicate, categorise, summarise, and respond to substantive comments, and that is human work performed by people with other duties. Regulations.gov, the federal eRulemaking portal operated by the General Services Administration, aggregates dockets across well over a hundred federal agencies and hosts comment volumes in the tens of millions. Every fabricated submission consumes some fraction of a limited pool of attention that was supposed to be spent reading what actual members of the public said.

The asymmetry here is brutal and it is the same asymmetry that closed a well known open source bug bounty programme, which we covered in the piece on bug bounty triage. Producing a plausible submission is now nearly free. Evaluating one is not. Any open intake system with that cost structure eventually collapses or closes.

Harm three: the legitimacy cost, which nobody prices

A citizen who learns that a federal docket contained 18 million fabricated comments, some in the names of dead people, draws a reasonable conclusion about whether their own comment mattered. That conclusion is corrosive and it is not recoverable through better process documentation. Participation is a habit, and habits die.

Why did the old defences stop working?

Agencies were not naive. After 2017 there was real effort, and the controls deployed were the sensible ones available at the time. Each of them has a specific failure mode now.

Duplicate detection, defeated by paraphrase

The workhorse control was textual similarity. Campaigns generate comments from a template, so identical or near identical text clusters together, and an agency can treat a cluster as one position held by many signatories rather than as many independent arguments. This worked well and it was honest, because organised campaigns are legitimate political activity and grouping them is not a judgment against them.

Generative text ended this. A language model will produce ten thousand comments arguing the same position in ten thousand distinct voices, with different sentence structures, different anecdotes, different levels of formality, and different apparent regional idiom. Similarity clustering sees ten thousand independent citizens. This is not speculation about the future, it is a property of the tools as they exist, and it is the single most important technical fact in this entire subject.

CAPTCHA, solved commercially

Challenge puzzles at the point of submission were the standard bot control. They are solved as a service for a fraction of a cent per challenge, and increasingly by the same models writing the comments. A CAPTCHA is a cost, and the cost is now below the value of a submission to anyone motivated enough to be doing this.

Email confirmation, which confirms an email

Requiring a working email address establishes that somebody controls an inbox. Inboxes are free and unlimited. This control was never strong and is now purely ceremonial.

Forensics, which arrives years later

The New York Attorney General's findings were published in 2021 about a docket that closed in 2017. Post hoc forensic analysis is genuinely valuable, and it is how we know what happened, but it operates on a timescale of years while the rule it concerns took effect on a timescale of months. Forensics is an autopsy, not a treatment.

And now, filing agents

One more change deserves naming. Agentic browsers can complete web forms on a user's behalf, which means the act of filing a comment can now be delegated to software by an ordinary person with no technical skill and no bad intent. This is not necessarily bad, and we will return to it, but it does mean the population of submissions now includes a category that did not previously exist: comments genuinely authorised by a real person who did not personally type them.

Why is identity the wrong answer here?

The obvious response to impersonation is to require everyone to prove who they are before commenting. It is obvious, it would work, and it would be a serious mistake. This section is the one that matters most, so it gets the space.

The constitutional position

The right to petition the government for redress of grievances is enumerated in the First Amendment, alongside speech, press, assembly and religion. Public comment is one of its most direct modern expressions.

Anonymous political speech carries its own strong protection. In Talley v. California (1960) the Supreme Court struck down an ordinance requiring handbills to identify their authors. In McIntyre v. Ohio Elections Commission (1995) the Court struck down a prohibition on distributing anonymous campaign literature, and its reasoning is worth understanding rather than merely citing: anonymity has an honourable tradition in American political advocacy, and the decision to remain anonymous may be motivated by fear of retaliation, concern about social ostracism, or simply a desire to have the idea judged on its merits rather than its author.

Now apply that reasoning to a comment docket. Consider who comments on federal rules. An employee of a regulated firm describing what actually happens inside it. A physician commenting on a payment rule their employer has an interest in. An immigrant commenting on immigration procedure. A person on disability benefits commenting on eligibility criteria. A local official disagreeing with their own administration. These are exactly the commenters whose input an agency most needs and who are most exposed if participation carries a name.

The chilling effect is not hypothetical

Any identity requirement produces a predictable sorting. The people deterred are not the well resourced campaigns, who will comply without difficulty and who employ counsel. The people deterred are individuals with something to lose, which is the same pattern documented for real name policies on platforms, discussed in the piece on verified human comments. An identity requirement on public comment would improve the docket's provenance while degrading its representativeness, and it is not obvious that this is a good trade.

The federated identity option, examined honestly

The United States has a government identity service used across federal agencies for authenticated interactions. It is well built and it solves a real problem for services like benefits applications, where the agency must know exactly who it is dealing with. Routing public comment through it would substantially reduce impersonation.

It would also collect identity for an activity where identity is not required, tie every comment to a verified federal account, and be unavailable to non citizens, to people abroad who are affected by American rules, and to anyone without the documentation such systems require. Comment periods regularly draw international participation on rules with extraterritorial effect, and a citizens only channel would be a substantive narrowing of who gets heard.

The general point, which applies well beyond this example: an identity system is the correct tool when the agency's decision depends on who you are, and the wrong tool when it depends only on whether you are a person.

What does a rulemaking actually need to know?

Strip the problem down and the agency needs remarkably little.

It does not need to know the commenter's name. It never did. A comment is weighed on the substance of what it says, and an anonymous comment raising a fatal technical flaw in a proposed rule is worth more than a signed one saying "I agree."

What it needs is narrower and quite specific. First, that a submission bearing a name was actually made by the person named, because a false attribution is a distinct wrong done to a specific individual, independent of any effect on the rule. Second, that a submission counted as one person's view represents one person, so that the docket's apparent composition is not a fabrication. Third, that organised and automated submissions are visible as such, because grouping a campaign is legitimate and mistaking a campaign for spontaneous public sentiment is not.

None of those three requires a name. All three are properties of the submission event rather than of the submitter.

What the agency wantsDoes it require identity?What actually establishes it
This comment's substance is worth consideringNoReading it. Anonymity is irrelevant.
The named person really filed thisNo, only a binding to the claimA signature from a device the named person enrolled
This is one person, not one operator with 40,000 accountsNoA uniqueness proof carrying no name
This block is an organised campaignNoDisclosure, plus the absence of independent proofs
This commenter is a licensed radiologist in OhioYes, when relevance depends on itA credential, voluntarily presented

Only the last row needs identity, and only because the commenter chose to make a claim about themselves that they want the agency to weigh.

How would a signed comment actually work?

The mechanism is unremarkable, which is the point. A comment portal already collects a submission. It additionally collects a signature over that submission, produced on a device the commenter controls.

Concretely: the commenter completes a short one time enrolment on their own phone, which runs a liveness check locally and derives a one way key. The face image never leaves the device and no template is stored anywhere. When they file a comment, the portal computes a hash of the submission and asks the commenter's device to sign it. What the portal receives back is a receipt.

{
  "docket":        "AGENCY-2026-0142",
  "comment_sha256": "9f2c...a71e",
  "submitted_at":  "2026-09-27T14:22:08Z",
  "attribution":   "named",
  "name_claimed":  "Priya Raman",
  "human_proof":   "unique-per-docket",
  "assistance":    "none",
  "sig":           "ed25519:5d81...c003",
  "key_id":        "manav-2026-09"
}

Two fields carry the weight. The human_proof field asserts that a distinct person produced this receipt within this docket, without asserting who. The attribution field records whether the commenter chose to attach a name, and if so, that the signing device is the one bound to that claimed name.

The receipt verifies offline. Anyone holding the published key can check it: the agency, a court reviewing the record, a journalist auditing the docket, a researcher three years later. Nobody has to call us, and nothing about the process depends on our continued existence, which is a property worth insisting on for anything that becomes part of a public record. The piece on offline verification covers why that independence matters more than it sounds.

The anonymous path is a first class citizen

This is the design decision that determines whether the whole thing is acceptable. A commenter can file with attribution: "anonymous" and no name field at all, and still carry a human_proof. The agency learns that a distinct person filed, and nothing else. That combination, a real human whose name nobody knows, is precisely what McIntyre protects and precisely what the current system cannot offer, because today an anonymous comment and a fabricated comment are indistinguishable.

Notice what this does. It makes anonymity stronger, not weaker. Right now, agencies discount anonymous comments partly because they cannot tell them apart from manufactured ones. Give anonymous comments a credible humanity proof and they become weightier than they are today, while revealing less than a name.

Delegated filing, handled honestly

The assistance field exists because a rule that pretends nobody uses tools would be ignored. A person may legitimately ask software to draft or file on their behalf. The useful design records that fact rather than prohibiting it: a delegated filing carries a signature from the human who authorised it, plus a disclosure that an agent submitted it, using the same scoped delegation model described in the piece on delegation chains. An agency can then treat "human wrote and filed" and "human authorised, agent filed" differently if it wishes, which is a policy choice made with information rather than a guess made without it.

Weight the docket, do not police the door

Here is the recommendation, and the reasoning behind it is constitutional rather than technical.

A system that blocks unproven comments creates a gate, and a gate on petitioning the government is a serious thing that invites serious challenge. A system that annotates comments and lets the agency weigh them creates no gate at all. Everyone may still file. Nobody is turned away. The agency simply knows more about the composition of what it received, and can say so publicly.

An agency publishing a docket summary could report: 402,118 comments received, of which 71,204 carry a unique human proof, 9,890 carry both a proof and a verified name attribution, 240,000 are identified as part of three disclosed organised campaigns, and the remainder are unproven. That single paragraph would have made the 2017 episode impossible to misrepresent, and it excludes nobody.

ProposalStops impersonation?Cost to civil libertiesAssessment
Do nothing, forensics afterwardsNoNoneCurrent state. Harm already demonstrated.
CAPTCHA at submissionNoMinimal, some accessibility costCeremonial. Solved commercially.
Duplicate text clusteringNoNoneWas good. Defeated by paraphrase.
AI text detectionNoModerate, false accusationsUnreliable and punishes non native writers.
Mandatory government identity loginYesSevere. Chills exactly the commenters most worth hearing, excludes non citizens.Effective and constitutionally hazardous.
Mandatory human proof, no identityMostlyModerate. Still a gate, still excludes the unenrolled.Better, but a gate is a gate.
Optional human proof, published weightingReduces sharplyLow. Nobody is excluded from filing.Recommended. Information rather than permission.

What this cannot do

The honest limits, and there are several substantial ones.

It does not stop paid real humans. An organisation that pays ten thousand actual people to file ten thousand actual comments produces ten thousand valid proofs. That is astroturf, it is arguably protected activity, and no identity primitive touches it. What changes is the price. Fabricated comments cost approximately nothing; recruited humans cost real money, which is a meaningful constraint even if it is not a wall.

It does not judge whether a human wrote the words. A person who asks a model to draft their comment, reads it, agrees with it, and signs it has made a genuine submission, and treating that as fraud would be both wrong and unenforceable. The proof establishes that a person stood behind the filing. It says nothing about composition, and it should not.

Uniqueness across dockets is harder than uniqueness within one. Preventing one person from filing once in each of forty dockets is straightforward. Doing so without creating a cross docket identifier that tracks a citizen's political participation across agencies is a genuinely difficult privacy problem, and the cryptographic technique for it, nullifiers, is roadmap work for us rather than something we ship today. Anyone claiming this is solved should be asked to show the construction. Until it is, per docket uniqueness is what is honestly available, and a linkable cross docket identifier would be worse than the disease.

Enrolment friction is real. Most people comment on a federal rule once in their lives, if ever, usually because something has directly affected them. Asking that person to enrol on a phone before speaking is a burden, and if the enrolment path is the only path, the system has re-created the gate it was meant to avoid. This is why the optional design matters so much.

Not everyone has a smartphone. The same limitation that applies everywhere in this domain applies here, and it lands hardest on elderly, rural and low income commenters. Any deployment needs a staffed alternative treated as a normal route rather than an exception.

None of this is a legal opinion. Whether a particular scheme survives First Amendment scrutiny is a question for courts and for counsel, not for a technology vendor, and anyone telling you their product is constitutional is selling something.

What to do this week

For agency staff, portal operators, and anyone running a consultation or petition system.

  1. Separate your categories. Falsely attributed, mass campaign, and machine generated are three different things needing three different responses. Adopt the distinction used in the Administrative Conference's Recommendation 2021-1 rather than lumping everything as spam.
  2. Measure your current exposure. Take a recent closed docket and ask what fraction of comments you could presently defend as coming from distinct real people. The answer is usually zero, and knowing that is the start.
  3. Publish composition, not just totals. Even before any new control exists, reporting how many submissions came through bulk upload versus the web form tells the public something true and costs nothing.
  4. Create a false attribution reporting route. Someone who finds a comment in their name should be able to say so in one step, and that report should attach to the record. Many portals have no such route at all.
  5. Stop treating anonymous as suspicious. Anonymity is protected. Conflating it with fraud is the reflex that pushes systems toward identity requirements.
  6. Pilot optional proof on one docket. Make it voluntary, publish the resulting composition, and see what proportion of commenters opt in. That single number will tell you more about feasibility than any amount of internal debate.
  7. Ask vendors what they retain. Any humanity proof that requires storing a biometric template or a document image has moved the risk rather than removed it. The question to ask is what survives the check.

You can see the shape of the gate in the humans first demo, and the mechanics of the receipts in the developer documentation.

Frequently asked questions

How many of the FCC net neutrality comments were fake? The New York Attorney General's 2021 investigation found that nearly 18 million of the more than 22 million comments filed in the 2017 proceeding were fabricated, including millions submitted under the identities of real people without their consent. Campaigns on both sides of the policy question contributed to the total.

Can AI flood a public comment period? Yes, and the important change is qualitative rather than quantitative. Agencies previously grouped campaign comments by textual similarity. Generative text produces large volumes of individually distinct comments, which defeats similarity clustering and makes an organised campaign look like spontaneous public sentiment.

How can agencies verify that commenters are real people? By accepting a submission signature produced on a device the commenter enrolled, which asserts that a distinct person filed without revealing who. The agency verifies the receipt against a published key, offline, and can publish how many submissions carried one. No name, document or biometric is collected.

Does verifying commenters violate the First Amendment? Requiring identity to petition the government raises serious constitutional questions, given the protection anonymous political speech has received in cases such as Talley v. California and McIntyre v. Ohio Elections Commission. A voluntary proof carrying no identity, used to weight rather than to exclude, is a materially different proposition. This is not legal advice.

Does comment volume actually affect a rule? Not directly. Notice and comment rulemaking under the Administrative Procedure Act requires agencies to consider significant comments and give reasoned explanations, not to count votes. The harms from fabricated comments are the contaminated record, the wasted processing, and the damage to public confidence.

What about comments filed by an AI agent for a real person? That should be permitted and disclosed rather than banned. A delegated filing can carry a signature from the human who authorised it plus a flag that an agent submitted it, so the agency can apply whatever policy it chooses with the facts in front of it.

Would this stop organised advocacy campaigns? No, and it should not try. Organised campaigns are legitimate political participation. The aim is to make them visible as campaigns rather than allowing them to be presented as spontaneous individual sentiment, and to make fabricated participation expensive.

Sources

  1. Office of the New York State Attorney General, 2021 report on fake comments in the FCC net neutrality proceeding, including findings on falsely attributed submissions: ag.ny.gov
  2. Administrative Conference of the United States, Recommendation 2021-1, Managing Mass, Computer-Generated, and Falsely Attributed Comments: acus.gov
  3. Federal Communications Commission, Electronic Comment Filing System, the public docket system used in the proceeding: fcc.gov/ecfs
  4. Regulations.gov, the federal eRulemaking portal operated by the General Services Administration: regulations.gov
  5. Administrative Procedure Act, notice and comment rulemaking, 5 U.S.C. 553: law.cornell.edu
  6. McIntyre v. Ohio Elections Commission, 514 U.S. 334 (1995), on protection for anonymous political speech: supreme.justia.com
  7. Talley v. California, 362 U.S. 60 (1960), striking down a compelled authorship identification ordinance: supreme.justia.com
  8. US Government Accountability Office, reporting on federal rulemaking and comment processes: gao.gov
A docket does not need to know your name. It needs to know that the name on a comment was put there by the person it belongs to, and that one person left one comment.