Designing a benchmark that is not a vendor opinion poll
The industry has no shortage of statistics about authorisation controls. It has a shortage of statistics anyone should act on, and the difference is entirely in the methodology section nobody publishes.
What makes a security statistic worth citing?
Four things, all of which most published figures lack: a sampling frame you can inspect, behavioural questions rather than self-assessment, a disclosed sponsor, and published raw data. A number without those is marketing that has learned to look like research.
- Sponsored surveys with unpublished sampling frames produce numbers that cannot be interpreted, however large the sample.
- Self-assessment questions measure belief; behavioural questions about specific recent events measure something closer to practice.
- Publishing the instrument, the frame and the raw distributions is what makes a benchmark checkable, and almost nobody does it.
Part of Research, datasets and methods
Why most industry statistics are unusable
| Problem | Effect |
|---|---|
| Sponsor has a product in the category | Question wording tends toward the sponsor's framing |
| Sampling frame unpublished | You cannot tell who was asked or who declined |
| Self-assessment questions | Measures belief about practice, not practice |
| Response rate unreported | Non-response bias is unknowable |
| Only aggregates published | Distributions and outliers are hidden |
| Year-on-year comparisons across changed instruments | Trends are artefacts of rewording |
Any one of these makes a figure hard to interpret. Most published surveys have several, which is why practitioners treat them as marketing and cite them anyway for lack of alternatives.
Behavioural questions instead of assessments
The single highest-leverage design choice. Ask about specific recent events rather than about general practice.
# Self-assessment — measures belief
"Do you require out-of-band verification for payment
changes above $10,000?" → 94% yes
# Behavioural — measures practice
"Think of the most recent payment instruction change
your organisation processed. Was verification
performed through a channel other than the one that
delivered the request?"
→ yes / no / do not know / no such change in 90 days
# The second produces a much lower number, and a
# large 'do not know' category that is itself a finding.
The do-not-know category is usually suppressed in published surveys because it looks like a defect in the instrument. It is a finding: an organisation that cannot say whether its control operated has an evidence problem regardless of the control's design.
What a defensible design looks like
- Publish the frame. Who was eligible, how they were identified, how many were approached, how many responded.
- Blind the sponsor. Respondents should not know who commissioned it while answering.
- Ask behaviourally. Specific recent events, with an explicit do-not-know option.
- Publish the instrument verbatim. Every question, in order, with the response options.
- Publish distributions, not only means. A bimodal distribution reported as an average is misleading.
- Keep the instrument stable across years, and flag every change.
Point six is what makes a benchmark valuable over time and what most publishers sacrifice, because refreshing the questions makes each year's report feel new.
What is worth measuring here
Four things, all behavioural, all answerable from an organisation's own records if they exist.
| Measure | Why it matters |
|---|---|
| Proportion of in-scope payments with a documented verification | Actual control coverage rather than policy |
| Proportion of respondents who cannot determine that | An evidence gap, independent of the control |
| Time from fraudulent payment to detection | Determines recovery odds |
| Whether the organisation has attempted a verification bypass test | Distinguishes assumed controls from tested ones |
The second row is the one no vendor-sponsored survey reports, and it is arguably the most important number in the field.
Stating the conflict
A benchmark published by a party with a product in the category has a conflict, and pretending otherwise is not credible.
The honest handling is to say so prominently, publish everything needed for someone to check the work, and accept that findings unfavourable to the sponsor's position must be published too. A benchmark whose results always support the publisher's product is not a benchmark.
What a reader should demand
- The sampling frame and response rate, not just the sample size
- The instrument, verbatim
- The raw distributions
- A statement of who paid for it
- An explicit list of what the data does not support concluding
The last is the clearest signal of good faith, and it is the one item almost never present.
The instrument, in full
An instrument worth trusting is short, behavioural and published before fieldwork, so nobody can claim afterwards that a question was added once the data looked interesting. This is the core set.
| # | Question | Type |
|---|---|---|
| 1 | How many payment releases above your approval threshold occurred last month? | Count |
| 2 | Of those, how many required a cryptographic signature from a named person? | Count |
| 3 | How many were approved via chat, email or a screenshot? | Count |
| 4 | How many production deployments last month were executed by an autonomous agent? | Count |
| 5 | Of those, how many were gated by a human approval at the endpoint? | Count |
| 6 | Has your organisation experienced an approval that was later disputed as unauthorised? | Yes / No / Prefer not to say |
| 7 | If yes, what evidence was produced to resolve it? | Categorical |
| 8 | How long, in minutes, from a decision to revoke an agent to the agent stopping? | Duration |
Note what is absent: no question asks the respondent to rate anything. Every answer is a count, a duration or a category, and every one of them exists in a system the respondent can check.
What a defensible design states up front
| Item | Why |
|---|---|
| Sampling frame and recruitment | Panel composition drives the result more than the questions do |
| Response rate and non-response handling | An unstated response rate usually means a poor one |
| The instrument verbatim | So nobody can add a question after seeing the data |
| Sponsor and conflict | The sponsor sells something; say so plainly |
| Raw data release | So a sceptic can reanalyse and disagree publicly |
Objections and honest limits
“A vendor cannot run a credible benchmark.” It can, if the conflict is stated, the instrument is published in advance and the raw data is released. Those three make the result checkable regardless of who paid for it.
“Behavioural questions get lower response rates.” They do, because they require looking something up. That is a cost worth paying, and the honest thing is to publish the response rate rather than compensate with easier questions.
What a reader should demand
- The sampling frame, described concretely. Not “senior decision-makers”.
- The response rate. Absent usually means low.
- The instrument verbatim. Question wording drives answers.
- The sponsor, stated plainly. Everyone has one.
- The raw data. So a sceptic can disagree in public.
Terms used here
- Sampling frame
- The population a survey drew from, which shapes the result more than the questions do.
- Behavioural question
- One asking what happened, answerable from records, rather than asking for a judgement.
- Self-assessment bias
- The tendency for maturity ratings to track self-image rather than practice.
Frequently asked questions
What is the single biggest design flaw in industry surveys? Self-assessment questions. They measure what respondents believe about their practice rather than what their records would show.
Why publish the do-not-know rate? An organisation that cannot say whether its control operated has an evidence problem regardless of the control's design. It is a finding, not an instrument defect.
Why does instrument stability matter? Year-on-year trends are meaningless if the questions changed. Refreshing questions makes each report feel new and destroys comparability.
Can a vendor publish a credible benchmark? Only by disclosing the conflict, publishing everything needed to check the work, and publishing findings unfavourable to their own position.
Why are most security statistics unusable? Opaque sampling, self-assessed answers, undisclosed sponsorship and no raw data. Any one of those makes a figure uncitable.
Can a sponsored benchmark be credible? Yes, if the conflict is stated, the instrument is published before fieldwork and the raw data is released.
Why publish the instrument in advance? So nobody can add or reword a question after seeing which answers look good.
Where this fits in Manav
The questions above measure whether approvals produce evidence. Manav is what turns a “no” in question two into a “yes”.
Sources and further reading
- Uniform Electronic Transactions Act (ULC)
- AAPOR Standard Definitions and disclosure standards
- Published critiques of vendor-sponsored security research.
- FBI IC3 2025 Internet Crime Report