Data

Who volunteers

Men who agree to be measured for a study are not a random draw, and the direction of that selection is a standing caveat on every reported mean.

4 min readData

No measured study recruits its participants by lottery. Every man in a size study made a choice to take part, and that choice is a selection filter sitting between the true population and the number the study eventually reports.

The mechanism

Volunteer bias happens whenever participation is optional and the decision to participate correlates with the thing being measured. If men who believe themselves above average are more willing to volunteer for a measurement study than men who do not, the resulting sample skews upward regardless of how carefully the actual measuring is done. If the opposite is true - if men with concerns avoid being measured, or men indifferent to the subject volunteer at the same rate as anyone else - the skew runs the other way, or vanishes. The mechanism does not by itself say which direction wins; it says a filter exists. It sits alongside, but is distinct from, the self-report inflation this site covers elsewhere - that is a bias in how a man reports his own figure, while volunteer bias is a bias in who ends up in the room to be measured at all, and a study can suffer from either, both, or neither. Volunteer bias inside a single study is also a different mechanism from publication bias across the literature as a whole, which is about which finished studies make it into print rather than who walks into any one of them.

Why this is a standing caveat, not a fixable error

Careful measurement technique fixes reading error. It does nothing about who is standing in front of the ruler. This is why volunteer bias travels with a reported mean as a caveat rather than as a correction that can be applied after the fact - the paper reporting 13.12 cm for erect length cannot tell you how that figure would shift if the volunteering process worked differently, only that it might. Veale et al. (2015), writing in BJU International, is explicit that this kind of selection is a limitation inherent to any study of this subject, not one specific to a particular sample.

How study design pushes back on it

Two design choices reduce, without eliminating, this problem.

Consecutive clinic attenders. Recruiting every eligible man who passes through a clinic in a given period, rather than posting a call for volunteers, removes the step where a man decides "this study is for people like me" before agreeing to take part. It does not remove the earlier selection of who attends that clinic in the first place, but it removes one additional layer of self-selection on top of it.

Recruitment framed around something other than size. A study recruiting for a general urology or health assessment, where the size measurement is one item on a longer form rather than the advertised purpose, draws participants who are not choosing in based on how they feel about their own measurement. This is a materially different recruitment pool than a study advertised as being about size specifically, which selects for men already thinking about the subject before they walk in.

Why the bias cannot be measured after the fact

Correcting for volunteer bias statistically requires knowing something about the men who declined to take part, and that group is, by definition, absent from the data - there is no record to compare the volunteers against. This is different from a measurement error, which repeat readings and known instrument limits can bound with a number; volunteer bias can only be reasoned about directionally, from what is known about recruitment method, not calculated from the dataset itself. It is one of the reasons a single large pooled study is not automatically more trustworthy than a smaller one with a cleaner recruitment method - sample size fixes precision around whatever population actually got measured, and does nothing to fix who that population was.

What this means for reading the pooled figure

The direction and size of any resulting skew in Veale 2015 is not something this site can state as a number - the paper does not quantify it, and inventing a magnitude would be exactly the kind of statistic this site does not print. What can be said directionally: the pooled mean is a statement about men willing to be measured under the recruitment method each contributing study used, and that population is not guaranteed to match a purely random draw of adult men, a caveat that sits alongside every mean and every nomogram this literature produces.

This is also a reminder of how different a measured figure is from a subjective one. A rating on Rate Cock carries its own selection effects - who submits a photo - but they are a different kind of bias attached to a different kind of number. An AI model scoring an uploaded image inherits whoever chooses to upload, a comparative score built from submitted photos inherits the same kind of filter, and a human judge only ever sees what gets submitted for review - volunteer bias is not unique to clinical measurement, but the direction and mechanism differ enough between a ruler study and a photo platform that neither should be assumed to correct for the other.

Read next

Full archive