Data

Why the numbers differ between papers

Two clinician-measured studies can report means a centimetre apart because of state, method, population and instrument, before anatomy is involved at all.

By 3 min readData

Guides on Data: How a size study is built, Reading a nomogram without fooling yourself, The Veale meta-analysis, read properly

Two studies, both using trained health professionals and a ruler, can report means that sit a centimetre apart. That gap does not require any real difference between the men measured in each - it can be produced entirely by differences in how the two studies were run. Veale et al. (2015) reports substantial heterogeneity across the studies it pooled, which is the formal statistical signal that the individual studies are not just noisy versions of the same underlying number - they are measuring somewhat different things and calling the result by the same name.

The sources, and which way each pushes

State measured. A study reporting flaccid length will report a lower, more variable mean than one reporting erect or stretched length, for reasons that have nothing to do with the population - flaccid length itself varies with temperature and recent activity, so a flaccid-only study's mean depends partly on when in the day the measuring happened.

Bone-pressed or not. A study that presses the ruler to the pubic bone removes the pre-pubic fat pad from the reading; a study that does not press to bone includes it, and the fat pad varies enough between men that this single protocol choice moves a study's mean by a meaningful amount, independent of anything anatomical. How consistently studies actually do this is its own separate question worth reading if you want the detail.

Posture. Standing versus lying down changes the fat pad's position and the shaft's resting angle, which shifts a non-bone-pressed reading more than a bone-pressed one.

How the erection was obtained. The literature draws erect measurements from a few different sources - self-stimulation in private, pharmacologically induced erection in a clinical setting, or self-report of an at-home reading - and these are not guaranteed to produce the same rigidity or the same reading conditions.

Sample. A clinic-recruited sample, a general population sample, and a sample recruited specifically for a size study are three different populations before a ruler is involved. Who volunteers for this kind of study is not a random draw, and which pool a paper drew from moves its mean.

Instrument and observer. A rigid ruler and a flexible tape used for length do not agree, and one observer's technique does not perfectly match another's even with a shared written protocol.

Why this matters more than picking "the right number"

None of these differences are errors in the ordinary sense - a study that measures flaccid length and says so is not wrong, it is answering a different question than a study that measures bone-pressed erect length. The mistake happens downstream, when someone lines up means from several papers as though they were repeated measurements of one fixed quantity and treats the spread between them as sampling noise. It is not sampling noise; a meaningful part of it is protocol.

Two of these sources are large enough, and specific enough, to deserve their own treatment rather than a line here: how stretched-length figures in particular vary between protocols because tension is hard to standardise in words, and the statistical shape of disagreement between pooled studies, which is what a meta-analysis is reporting when it flags heterogeneity rather than just publishing a single combined mean.

Reading a table of study means

The practical habit this suggests: never read a table of study means without also reading the methods column. A mean without its state, its bone-pressed status, its sample, and its instrument is not comparable to the mean in the next row, whatever the table's formatting implies by putting them side by side.

This is a different kind of disagreement from the one that shows up when self-reported figures diverge from measured ones - here every study in question used a trained measurer, and the gaps still appear. It is also a different question from a rating disagreeing with another rating; a score from a judged platform or an AI-derived estimate can vary between reviewers or models for reasons that have nothing to do with study protocol, and comparing that kind of variation to protocol-driven heterogeneity in clinical literature mixes two unrelated sources of disagreement. If the number you actually want is a subjective judgement rather than a defensible mean, a rating service is answering a different question than any of the papers discussed here, and its internal consistency is a separate matter from study heterogeneity - as is a distribution of scores drawn from photographs rather than a ruler.

Read next

Full archive