Manav.id
Platforms ยท 18 min read

Can a comment section prove its commenters are human without knowing who they are?

Every platform that has tried to solve bots has offered users the same deal: give us your identity, or share the room with software. That choice is a false one, and it exists because nobody built the primitive that sits between the two. Humanity, uniqueness and identity are three separate properties, and a comment section only needs the first two.

A woman in a country where it is unwise to say certain things out loud has been posting for six years under a name that is not hers. She writes carefully. She has never been abusive, never been reported, never been suspended. Last month the platform asked her to verify her account by photographing a government identity document, because an automated system had flagged her posting pattern as inauthentic.

She will not do it, and she is right not to. The document would tie the account to her legal name in a database she has no visibility into, held by a company in a jurisdiction that may honour a request from hers, and retained for a period nobody has told her. The rational move is to abandon six years of writing and start again. So that is what she does.

The same week, on the same platform, an operator running four hundred accounts through a language model and a pool of rented phone numbers passes every check without incident, because the checks were never designed to catch someone who is willing to pay a small amount of money.

Both of those outcomes come from the same mistake. The platform reached for identity when what it actually needed to know was something much narrower, and much less dangerous to hold.

Short answer: Yes. A comment section needs to know two things: that a participant is a person, and that one person is not fifty accounts. Neither requires a name. A device bound proof can establish both while storing only a one way key, with the biometric check happening on the phone and no document, image or template retained by anyone. Identity is a third, separate property that most platforms request by reflex.

What choice do platforms actually offer today?

Two options, both bad, and a great deal of engineering effort spent making each one slightly less bad.

Option one: detection

The platform keeps accounts cheap and tries to work out which ones are not people, using signals like posting cadence, device characteristics, network origin, linguistic patterns and graph position. This is the dominant approach and it has a well understood failure mode, which is that it produces two kinds of error and only one of them is visible to the platform.

The invisible error is the operator who passes. The visible error is the person who does not, and that error has a specific texture that anyone who spends time online will recognise: a real human being, often one with an unusual usage pattern, being told by an automated system that they are a robot, with an appeals process that is either absent or another automated system. Shift workers, people with disabilities using assistive technology, people on shared or institutional connections, people who post in bursts because that is when they have time. The false positive rate is not evenly distributed. It falls on people who are already at the edges of the usage distribution.

Independent measurement suggests automated traffic is now roughly half of all web requests (Imperva Bad Bot Report), and the detection task has become considerably harder since conversational models made the linguistic tells disappear. Posting cadence and writing style were the two strongest signals a decade ago. Both are now trivially imitable, and they were the signals that could be gathered without invading anyone.

Option two: identity

So platforms escalate to identity. Phone number verification, which stops nobody, because numbers are rented in bulk and the market for them is mature and cheap. Paid verification badges, which prove that somebody had a working payment card, a property that is also purchasable at scale and which conspicuously does not mean what users assumed a checkmark meant. And document verification, which does work, at the cost of building exactly the database that should not exist.

The document approach creates a permanent liability out of a momentary question. The platform wanted to know one bit of information, whether this is a person. It now holds a photograph of a passport, and it holds it for every user who complied, and that collection is a target for the rest of its existence. We covered this dynamic in detail in the piece on age verification, where the same pattern plays out under regulatory pressure and with the same outcome: users evade, the compliant minority take on the risk, and the stored data eventually escapes.

Regulation has increased the pressure without resolving the tension. The European Digital Services Act imposes systemic risk duties on large platforms, including for inauthentic behaviour and manipulation, and the UK Online Safety Act imposes its own set. Both create an obligation to act. Neither tells a platform how to act without collecting identity, because the primitive that would allow it was not in general use when the rules were drafted.

Meanwhile the harm the rules are meant to address keeps growing. The US Federal Trade Commission reported that consumers lost 2.1 billion dollars in 2025 to fraud that began on social media, an increase of roughly eight times since 2020 (FTC, 2026). Those losses do not come from clever technology. They come from accounts that a person believed belonged to another person.

Why is pseudonymity worth protecting?

Before proposing anything, it is worth being unambiguous about this, because a great deal of writing about bots treats anonymity as the problem and therefore treats identity as the cure. That framing is wrong, and it has caused real harm to real people.

Pseudonymity is not a loophole. It is load bearing infrastructure for a long list of people whose participation we should want.

It is the person discussing a diagnosis they have not told their employer about. The one in an abusive relationship who needs a community and cannot leave a trail. The employee describing conditions at a company that would fire them for it, which is precisely the speech that most needs to exist. The person in a country where the wrong opinion is a criminal matter. The teenager working out who they are before they are ready to tell anyone. The person who simply does not want a stranger reading their political opinion to be able to find out where their children go to school.

Every real name policy in the history of the internet has been introduced with the promise that it would improve civility, and the evidence for that promise has been consistently weak. What real name policies reliably produce is a shift in who participates. The people who leave are not the abusive ones, who are frequently happy to attach their names to what they say. The people who leave are the ones with something to lose.

So the design constraint is not negotiable, and it should be stated first rather than added at the end as a mitigation: a system that requires identity has failed before it starts. Not failed ethically, although that too. Failed at the job, because the users whose participation is most worth protecting will not use it, and the adversaries will.

Humanity, uniqueness, identity: three properties, not one

Here is the conceptual move that makes the whole problem tractable, and it is genuinely simple once stated. Platforms have been conflating three distinct properties, because historically the only tool available established all three at once.

PropertyThe question it answersWhat a comment section needs it forWhat it reveals
HumanityIs there a person here at all?Excluding automated participationNothing about which person
UniquenessIs this person the same one as that account?Stopping one operator running many accountsNothing about who they are
IdentityWhich specific named person is this?Almost nothing, for most communitiesEverything, permanently

Read the last column. The first two properties are cheap to hold and safe to lose. The third is the one that ends careers and occasionally lives, and it is the one platforms ask for.

The reason they ask for it is not malice. It is that a passport establishes all three properties in a single artifact, and for a long time no artifact established only the first two. When your only instrument is one that measures everything, you measure everything, and then you are responsible for storing it.

An analogy that holds up well: a nightclub door needs to know you are old enough. It does not need to know your address, your organ donor status, or your middle name, all of which are printed on the identity document you hand over. The bouncer reads one line and returns the card, and the system works because nothing is retained. Now imagine the club photocopied every document and kept the copies in a filing cabinet in the back office forever. That is what document verification online is, and the difference is not that the club is more trustworthy. The difference is that the club gives the card back.

What would a human proof actually prove?

The artifact is deliberately, almost disappointingly, small. That is the design goal.

A person enrols once. The check runs on their own device: an on device match with a liveness challenge that a photograph or a recorded video cannot satisfy. The biometric never leaves the phone. What is produced and retained is a one way key, a value derived from the check that cannot be reversed into a face, an image or a template. There is no document, no name, and no identifier that means anything outside this context.

What the platform receives looks like this:

{
  "claim":        "unique_human",
  "context":      "example-forum",
  "subject_key":  "b7d1...9e04",   // one-way, context-scoped
  "assurance":    "device_bound_liveness",
  "issued_at":    "2026-09-19T11:04:22Z",
  "expires_at":   "2027-09-19T11:04:22Z",
  "signature":    "ed25519:MEQCIF9..."
}

Three things about that object are worth dwelling on.

There is no identity field, because there is no identity. Not redacted, not encrypted, not held elsewhere under a policy. It was never collected. A subpoena served on the platform, or on Manav, produces nothing about who this person is, because nobody knows.

The key is scoped to the context. The value the forum sees is derived for that forum. The same person enrolling on a different platform produces a different value, so the two accounts cannot be correlated by comparing keys. This is what stops the proof becoming a tracking identifier, which would be a worse outcome than the problem it solves.

It verifies offline. A moderator, an auditor or a researcher can check the signature against a published key without calling Manav, without Manav learning that the check happened, and without the platform having to be trusted. Verification is a local computation:

from manav import verify

proof = request.headers["X-Human-Proof"]
ok, claim = verify(proof, published_key)          # no network call

if ok and claim["claim"] == "unique_human":
    allow_post(subject_key=claim["subject_key"])   # one key, one voice
else:
    queue_for_review()                             # not blocked, reviewed

Note the else branch. It does not reject. This matters, and we will come back to it.

Gate the action, not the reading

One implementation decision does most of the work in keeping this proportionate. The proof should gate the actions that carry abuse potential, which means posting, replying, direct messaging and voting, and it should never gate reading.

Reading is the overwhelming majority of use, it is where the value of an open web lives, and there is no abuse case that requires it to be gated. A design that puts a check between a person and information has become something other than an anti abuse control. Keep the reading open and put the proof where the cost of abuse actually is.

The two tier internet objection

This is the strongest objection to everything above and it should not be waved away, because the failure mode it describes is realistic and would be bad.

The objection runs like this. Introduce a verified human badge and you have created two classes of participant. At first the unverified tier is merely less visible: sorted lower, rate limited, shown with a warning. Then it becomes the default that nobody reads. Then platforms notice that verified users are more valuable to advertisers and the incentives point one way. Within a few years the unverified tier is functionally dead, and a proof that was introduced as optional has become mandatory through the accumulation of small product decisions, none of which was the decision to require identity.

That is a coherent prediction, and something like it has happened before with other optional signals. Anyone proposing a human proof needs to answer it, and the honest answer has several parts, none of which is completely satisfying.

The first part is that the proof deliberately carries no identity, so even a world where it becomes universal is materially different from a world where document verification becomes universal. If the tier gets enforced, what is enforced is that you had a phone and a face, not that you registered your name. That is a real distinction, and it is the reason to prefer this design over the alternative that is currently winning. It is not, however, a reason to be relaxed about enforcement.

The second part is that the design should make the unverified tier genuinely usable rather than nominally permitted, and that is a product commitment which can be broken. Rate limits rather than blocks. Review queues rather than rejection. Full reading access. These are choices a platform makes and can unmake.

The third part is the uncomfortable one. There is no technical mechanism that prevents a platform from making an optional signal effectively mandatory. None. That is a governance problem, and it will be answered by regulation, by public pressure, and by whether the people building these systems say clearly and repeatedly that a human proof must never become a precondition for speech. This document is one instance of saying it. It is not a guarantee, and anyone who tells you their protocol design guarantees it is selling something.

What can be said is that the alternative trajectory is worse and further along. Platforms are already escalating to document verification under regulatory pressure. The realistic choice is not between a human proof and an open internet. It is between a human proof and a passport photograph.

Honest limits

This is the section that matters most, because the domain punishes overclaiming and the failure modes here are social rather than technical.

It does not make discourse good. Not even slightly. Verified humans are perfectly capable of cruelty, dishonesty, bad faith and coordinated harassment, and some of the worst behaviour online comes from people posting under their real names with their employer in the bio. A human proof addresses automated participation and account multiplication. It has nothing to say about what people choose to do with a voice once they have one, and any pitch that implies otherwise is dishonest.

It does not stop coordinated campaigns by real people. Influence operations staffed by humans, brigading, and organised political messaging all pass this control completely, because everyone involved is a person. What changes is the cost structure. An operation that could be run by one person with a script now needs a payroll.

It can be defeated by renting humans, and that industry exists. Account farms that pay real people to complete verification steps are an established business in adjacent markets, and this control does not make them impossible. It makes them expensive per account rather than free at scale, and it converts an infinitely reproducible attack into one with a marginal cost. That is a meaningful change and it is not a solution.

Cross platform uniqueness is not shipped. The context scoped key that protects users from correlation also means one person can hold one account on each of many platforms, which is correct and desirable. Detecting that the same human holds fifty accounts across fifty platforms, without letting any of them correlate, requires nullifier constructions that Manav has not built. Describing that as available today would be false.

Enrollment is the trust bottleneck. Everything rests on the first check being a real person and a distinct one. Coordinated enrollment, where an operator moves through many devices with many cooperating humans, is the attack that matters, and defending it well at scale is unproven. It should be treated as an open problem rather than a solved one.

It excludes people without a suitable device. A proof that requires a modern smartphone excludes people who do not have one, which correlates with exactly the populations already least well served. Any deployment needs an alternative path, and if that path is a document upload then the design has quietly reintroduced the thing it was built to avoid. This is unresolved and it is the limitation that should worry a deployer most.

What to do this week

For anyone running a platform, a forum, a comment section or a community:

  1. Write down which properties you actually need. For each control you have, state whether it needs humanity, uniqueness, or identity. Most teams discover they have been collecting identity to answer a uniqueness question.
  2. Audit your false positive path. Find out what happens to a real user wrongly flagged as automated. If the answer is an automated appeal, you have a harm you are not measuring.
  3. Separate reading from acting in your gating logic. If any anti abuse control currently sits between a person and information, move it to sit before an action instead.
  4. Stop treating phone numbers as evidence of humanity. They are evidence of a small payment. Price your controls accordingly and be honest internally about what the signal is worth.
  5. Check what identity data you already hold and why. Documents collected for a verification flow that has since been replaced are a liability with no remaining purpose. Delete them.
  6. If you must add verification under regulatory pressure, ask vendors what they retain. The right answer to "what do you store" is a one way key. Any answer involving images or templates is a database you will be responsible for.
  7. Publish your policy on tiering. If you introduce any human signal, state publicly what unverified users can still do, and treat that as a commitment rather than a default. The commitment is the only real protection against the two tier outcome.

The humans first demo shows the enrollment flow and what the resulting proof contains, and the alternative to CAPTCHA demo shows the gate on an action. The developer documentation covers the widget modes and offline verification.

The general lesson

For thirty years the internet has treated identity as the answer to every trust question, because identity was the only instrument anyone had. Ask whether someone is real, and the system asks for a name. Ask whether one person is one account, and the system asks for a name. Ask whether someone is old enough, and the system asks for a name and a date of birth and a photograph of the document they are printed on.

This is the same error described across the identity failure map: reaching for a heavyweight property because the lightweight one was never built. The cost of that error is paid by the people with the most to lose from disclosure, and it is paid in silence, because the people who leave a platform rather than upload a passport do not file a support ticket.

A comment section does not need to know who you are. It needs to know that you are somebody, and that you are only one somebody. Those are small, answerable questions, and answering them properly is how a room full of strangers stays worth being in.

Frequently asked questions

How can I prove I am human online without uploading my ID? With a device bound proof. The check runs on your own phone with a liveness challenge, the biometric never leaves the device, and what is retained is a one way key that cannot be reversed into a face or a name. The platform receives a signed statement that a unique person is behind the account, and nothing else.

What percentage of social media accounts are bots? No trustworthy universal figure exists, and be suspicious of anyone who gives you one, since the platforms hold the data and have an incentive in how it is reported. For broader context, independent measurement puts automated traffic at roughly half of all web requests (Imperva Bad Bot Report), though web traffic and social accounts are different populations.

Is verified human the same as verified identity? No, and conflating them is the central mistake. Verified identity establishes which named person you are. Verified human establishes only that a person exists behind an account and that they are distinct from other accounts. The second can be proven while storing nothing that identifies anyone.

Is there an alternative to biometric personhood systems that keep a database? Yes. The distinction is where the check runs and what survives it. A system that captures a biometric to a central registry creates a permanent database. A device bound check runs on the user's phone and retains only a one way key, so there is no registry of faces to breach, subpoena or sell.

Will a human badge kill anonymity? It should not, because the proof carries no identity, so a verified pseudonymous account remains pseudonymous. The genuine risk is not disclosure but tiering: platforms gradually making unverified accounts unusable. That is a governance problem rather than a technical one, and no protocol design prevents it.

Does this stop trolls and harassment? No. It addresses automated accounts and one person running many accounts. People behave badly under their real names constantly. Anyone claiming a humanity proof improves the quality of discourse is overselling it, and the honest claim is narrower: it removes the free tier of manipulation.

What about users without a smartphone? This is the real limitation. Any deployment needs an alternative path, and if that path is document upload then the design has reintroduced the problem it was built to avoid. There is no fully satisfying answer today, which is why reading should never be gated and unverified participation should remain genuinely usable.

Sources

  1. Imperva Bad Bot Report, annual measurement of automated share of web traffic: imperva.com resource library
  2. US Federal Trade Commission, consumer protection data on fraud originating on social media: ftc.gov data spotlight
  3. European Commission, Digital Services Act, systemic risk assessment and mitigation duties (Articles 34 and 35): digital-strategy.ec.europa.eu
  4. Ofcom, Online Safety Act implementation and age assurance reporting: ofcom.org.uk online safety
  5. W3C Verifiable Credentials Data Model, for the standards context on holder presented claims: w3.org VC data model
A comment section does not need to know who you are. It needs to know that you are somebody, and that you are only one somebody.