Data

Exclusion criteria

Measured studies exclude men with conditions that affect the measurement, which is correct for their purpose and also trims the tails of the reported distribution.

4 min readData

Every measured study in the pooled literature excludes some fraction of the men who might otherwise take part. That is a normal and defensible feature of study design, not a flaw to be suspicious of. It is also a reason the reported distribution is narrower than the full population's would be, and worth knowing about before reading a mean as if it covers everyone.

Why exclusion exists at all

A study measuring length and girth by a standard protocol needs the protocol to apply cleanly to every participant, and the same physical circumstance that makes a reading unreliable for one man makes it unreliable for the study as a whole if included without note. Rather than take an unreliable reading and report it alongside reliable ones, a well-run study sets criteria in advance and excludes participants who would not fit them, then states the criteria in its methods section. This is standard practice across measurement research generally, not something specific or unusual to this literature. Height and weight studies, for instance, routinely exclude participants with conditions that would make a standing measurement unrepresentative, for exactly the same procedural reason.

Publishing the criteria alongside the results is what lets a later reader, or a later meta-analysis pooling several studies together, know what population a given mean actually describes. A study that measured everyone who walked through the door, without stating who was turned away and why, would be harder to compare to any other study, not easier - the criteria are part of what makes a figure interpretable at all, not an inconvenience layered on top of it.

The direction of the effect, without the detail

Exclusion criteria of this kind, by design, remove circumstances at the edges of what is anatomically typical, because those are the circumstances most likely to interfere with a clean, repeatable reading under the standard method. The mechanical consequence is that a study built this way is measuring a population that has already had some of its extremes trimmed before the mean and standard deviation are ever calculated. That does not make the reported figures wrong for what they claim to describe. It does mean the reported spread is narrower than the spread of an unfiltered population would be, and a reader placing an individual figure against a percentile chart built from excluded data should hold that in mind - a percentile is only as good as the reference population it was built from, and an excluded-tails population is a slightly different reference than an unfiltered one.

Why this belongs in the methods section, not a caveat

A study that did not exclude anything would not necessarily be more honest - it would likely be measuring some participants under conditions the protocol was not designed to handle, which introduces a different kind of error into the same dataset. The trade a study makes is: accept a slightly narrower reported population in exchange for every remaining reading being taken the same clean way. That trade is the right one for the study's actual purpose, which is describing what the standard method captures reliably, not enumerating every case an unfiltered population might contain.

Reading a mean with this in mind

None of this changes how to read your own number against the published figures - the clinical method itself is unaffected by what any given study excluded. It changes how confidently you should treat the tails of a distribution as complete. A mean and a standard deviation from a study with exclusion criteria describe the population that was actually measured, not a hypothetical unfiltered one, and the gap between those two populations is usually small but not zero. This is a separate question from who volunteers for a study in the first place, which shapes the sample before exclusion criteria are even applied. Selection happens twice, in other words: once when a man decides to volunteer, and again when the study's own criteria decide whether his measurement is usable, and the two filters are worth keeping distinct even though they push the same reported distribution in the same general direction, toward the anatomically typical middle of the range.

None of this touches subjective size judgements at all. Rate Cock and services built the same way sit entirely outside this question - a photograph-derived score, scored the way Rate Cock scores one, has no exclusion criteria in this sense, because it is not sampling from a measured population to begin with. The same is true of a human judge's verdict, a photo-based numeric score, or an AI model's estimate from an image: none of them draw from, or exclude from, a measured sample in the sense this post describes.

Read next

Full archive