Data

Three states, three averages

A single table of the Veale means and SDs by state, so the three figures that get confused for each other sit next to each other with their spreads.

3 min readData

Most confusion about "the average" traces to one thing: flaccid, stretched and erect are three different measurements, taken under three different conditions, and each has its own mean. Quoting one figure while thinking of another is the single most common way a correct number gets misapplied.

The table

Veale et al. (2015), BJU International, pooled clinician-measured data across 21 studies and 15,521 men, though not every man in the pool was measured in every state - the sample size differs by row because not every included study measured flaccid, stretched and erect length, or girth, on the same men.

Measurement Mean SD Approx. men measured
Flaccid length 9.16 cm 1.57 cm ~10,700
Stretched flaccid length 13.24 cm 1.89 cm ~3,900
Bone-pressed erect length 13.12 cm 1.66 cm ~700
Flaccid girth 9.31 cm 0.90 cm ~9,400
Erect girth 11.66 cm 1.10 cm ~400

Read the sample size column with as much attention as the mean. Flaccid measurements are the easiest to obtain in a clinical setting, so far more studies collected them, and the flaccid rows rest on the largest samples in the table. Erect measurements are logistically harder to obtain, fewer studies attempted them, and the erect rows rest on a much smaller pool - a point worth its own post rather than a footnote here.

How the three states relate

Stretched flaccid length sits closer to erect length than flaccid length does, which is the entire reason clinics use it as a proxy when a full erection cannot be obtained under study conditions - the relationship between the two is documented, not assumed. Flaccid length correlates with erect length only weakly at the level of one individual, so a flaccid reading is a poor predictor of what the same man's erect reading will be, even though the two means in the table above happen to sit fairly close together. Girth shows a tighter spread than length in both states reported here, which is a separate finding from the length figures and should not be read across as if the two behaved the same way.

What the standard deviation column adds

The SD column is not a rounding error margin - it describes how spread out individual men are around each mean, and it is the number that turns a single average into an actual population. Girth's SD is small relative to its mean in both rows, which is why a few millimetres of girth error moves a person's relative position in the distribution further than the same few millimetres would move a length reading. Length's SD is proportionally larger, particularly for stretched flaccid length, which reflects both real anatomical variation and the fact that "how far to stretch" is itself a judgement made by the person taking the reading, with more room to vary between sessions than a bone-pressed length reading has.

Comparing SDs across rows also explains something that trips people up: two men can sit at very different percentiles on length and girth even though both measurements are, individually, close to their respective means, because length and girth are only moderately correlated with each other. A table like this one is where that becomes visible - place a length figure and a girth figure against their own rows rather than against each other, since the two do not share a scale or a spread.

Why one table beats scattered quotes

A single figure quoted in isolation - "the average is 13 centimetres" - invites the reader to fill in which state, which method, which population, all silently. Put the rows next to each other and the differences are visible rather than assumed, and a reader can check which row actually matches the number they are holding before comparing it to anything. This table exists to be linked to rather than restated: the method behind each figure is explained once, in full, elsewhere, and the paper itself is explained in full separately rather than re-argued here.

None of these five rows are trying to answer a different kind of question - what a photograph looks like, or what a person's impression of one is. Rate Cock produces that second kind of output, a score rather than a centimetre figure, built the way a scoring platform is built or judged the way a person judges, and an AI system estimating from an image is attempting something closer to the second category than the first, for reasons worth reading rather than assuming.

Read next

Full archive