Manav.id
Compliance · 5 min read

Designing a benchmark that is not a vendor opinion poll

Designing a benchmark that is not a vendor opinion poll

The industry has no shortage of statistics about authorisation controls. It has a shortage of statistics anyone should act on, and the difference is entirely in the methodology section nobody publishes.

What makes a security statistic worth citing?

Four things, all of which most published figures lack: a sampling frame you can inspect, behavioural questions rather than self-assessment, a disclosed sponsor, and published raw data. A number without those is marketing that has learned to look like research.

Key takeaways
  • Sponsored surveys with unpublished sampling frames produce numbers that cannot be interpreted, however large the sample.
  • Self-assessment questions measure belief; behavioural questions about specific recent events measure something closer to practice.
  • Publishing the instrument, the frame and the raw distributions is what makes a benchmark checkable, and almost nobody does it.

Why most industry statistics are unusable

Self-assessment“How mature is your programme?”Answers track self-imageNot comparable between firmsUnfalsifiableProduces a flattering headlineBehavioural“How many wire approvals last month?”Answers track recordsComparableCheckable against realityProduces a boring, usable numbervs
ProblemEffect
Sponsor has a product in the categoryQuestion wording tends toward the sponsor's framing
Sampling frame unpublishedYou cannot tell who was asked or who declined
Self-assessment questionsMeasures belief about practice, not practice
Response rate unreportedNon-response bias is unknowable
Only aggregates publishedDistributions and outliers are hidden
Year-on-year comparisons across changed instrumentsTrends are artefacts of rewording

Any one of these makes a figure hard to interpret. Most published surveys have several, which is why practitioners treat them as marketing and cite them anyway for lack of alternatives.

Behavioural questions instead of assessments

The single highest-leverage design choice. Ask about specific recent events rather than about general practice.

# Self-assessment — measures belief
  "Do you require out-of-band verification for payment
   changes above $10,000?"                    → 94% yes

# Behavioural — measures practice
  "Think of the most recent payment instruction change
   your organisation processed. Was verification
   performed through a channel other than the one that
   delivered the request?"
  → yes / no / do not know / no such change in 90 days

# The second produces a much lower number, and a
# large 'do not know' category that is itself a finding.

The do-not-know category is usually suppressed in published surveys because it looks like a defect in the instrument. It is a finding: an organisation that cannot say whether its control operated has an evidence problem regardless of the control's design.

What a defensible design looks like

  1. Publish the frame. Who was eligible, how they were identified, how many were approached, how many responded.
  2. Blind the sponsor. Respondents should not know who commissioned it while answering.
  3. Ask behaviourally. Specific recent events, with an explicit do-not-know option.
  4. Publish the instrument verbatim. Every question, in order, with the response options.
  5. Publish distributions, not only means. A bimodal distribution reported as an average is misleading.
  6. Keep the instrument stable across years, and flag every change.

Point six is what makes a benchmark valuable over time and what most publishers sacrifice, because refreshing the questions makes each year's report feel new.

What is worth measuring here

Four things, all behavioural, all answerable from an organisation's own records if they exist.

MeasureWhy it matters
Proportion of in-scope payments with a documented verificationActual control coverage rather than policy
Proportion of respondents who cannot determine thatAn evidence gap, independent of the control
Time from fraudulent payment to detectionDetermines recovery odds
Whether the organisation has attempted a verification bypass testDistinguishes assumed controls from tested ones

The second row is the one no vendor-sponsored survey reports, and it is arguably the most important number in the field.

Stating the conflict

A benchmark published by a party with a product in the category has a conflict, and pretending otherwise is not credible.

The honest handling is to say so prominently, publish everything needed for someone to check the work, and accept that findings unfavourable to the sponsor's position must be published too. A benchmark whose results always support the publisher's product is not a benchmark.

What a reader should demand

The last is the clearest signal of good faith, and it is the one item almost never present.

The instrument, in full

An instrument worth trusting is short, behavioural and published before fieldwork, so nobody can claim afterwards that a question was added once the data looked interesting. This is the core set.

Core questions
#QuestionType
1How many payment releases above your approval threshold occurred last month?Count
2Of those, how many required a cryptographic signature from a named person?Count
3How many were approved via chat, email or a screenshot?Count
4How many production deployments last month were executed by an autonomous agent?Count
5Of those, how many were gated by a human approval at the endpoint?Count
6Has your organisation experienced an approval that was later disputed as unauthorised?Yes / No / Prefer not to say
7If yes, what evidence was produced to resolve it?Categorical
8How long, in minutes, from a decision to revoke an agent to the agent stopping?Duration

Note what is absent: no question asks the respondent to rate anything. Every answer is a count, a duration or a category, and every one of them exists in a system the respondent can check.

What a defensible design states up front

Published before fieldwork
ItemWhy
Sampling frame and recruitmentPanel composition drives the result more than the questions do
Response rate and non-response handlingAn unstated response rate usually means a poor one
The instrument verbatimSo nobody can add a question after seeing the data
Sponsor and conflictThe sponsor sells something; say so plainly
Raw data releaseSo a sceptic can reanalyse and disagree publicly

Objections and honest limits

“A vendor cannot run a credible benchmark.” It can, if the conflict is stated, the instrument is published in advance and the raw data is released. Those three make the result checkable regardless of who paid for it.

“Behavioural questions get lower response rates.” They do, because they require looking something up. That is a cost worth paying, and the honest thing is to publish the response rate rather than compensate with easier questions.

What a reader should demand

  1. The sampling frame, described concretely. Not “senior decision-makers”.
  2. The response rate. Absent usually means low.
  3. The instrument verbatim. Question wording drives answers.
  4. The sponsor, stated plainly. Everyone has one.
  5. The raw data. So a sceptic can disagree in public.

Terms used here

Sampling frame
The population a survey drew from, which shapes the result more than the questions do.
Behavioural question
One asking what happened, answerable from records, rather than asking for a judgement.
Self-assessment bias
The tendency for maturity ratings to track self-image rather than practice.

Frequently asked questions

What is the single biggest design flaw in industry surveys? Self-assessment questions. They measure what respondents believe about their practice rather than what their records would show.

Why publish the do-not-know rate? An organisation that cannot say whether its control operated has an evidence problem regardless of the control's design. It is a finding, not an instrument defect.

Why does instrument stability matter? Year-on-year trends are meaningless if the questions changed. Refreshing questions makes each report feel new and destroys comparability.

Can a vendor publish a credible benchmark? Only by disclosing the conflict, publishing everything needed to check the work, and publishing findings unfavourable to their own position.

Why are most security statistics unusable? Opaque sampling, self-assessed answers, undisclosed sponsorship and no raw data. Any one of those makes a figure uncitable.

Can a sponsored benchmark be credible? Yes, if the conflict is stated, the instrument is published before fieldwork and the raw data is released.

Why publish the instrument in advance? So nobody can add or reword a question after seeing which answers look good.

Where this fits in Manav

The questions above measure whether approvals produce evidence. Manav is what turns a “no” in question two into a “yes”.

See approval receipts →

Sources and further reading