Data
How a size study is built
Recruitment, exclusion, measurer, instrument, state and reporting: each design choice moves the reported figure, and knowing them is how you read any study.
A published mean is the end product of a long chain of decisions, and every link in that chain pushes the final figure in some direction. Reading a study well means reading past the headline number to the design choices behind it - and Veale et al. (2015), the meta-analysis this site treats as the standard reference, is a useful worked example because its own inclusion rule is explicit about several of these decisions.
Who gets recruited
The first decision is where participants come from, and it is rarely a random draw from the general population - that kind of sampling is expensive and impractical for a measurement study. Most measured studies recruit from urology or sexual-health clinics, from men already there for an unrelated reason, or from men who volunteer specifically because a study on this subject was advertised. Both of those recruitment paths introduce their own selection: a clinic population is not the general population by definition, and men who volunteer for a study about size are not a random subset of men who did not, in a direction that is not obviously predictable in advance. The clinic-versus-population gap gets its own full treatment, because it is the recruitment issue that shows up most often across the pooled studies behind Veale.
Who gets excluded
Recruitment brings people in; exclusion criteria then remove some of them, usually for good methodological reasons that still shift the reported mean. A study might exclude men with a diagnosed urological condition that would affect the measurement, men below a certain age, or men for whom a full erection could not be reliably obtained under the study's protocol. Each of those is defensible on its own terms - a condition that distorts the measurement should not be pooled with normal variation - but the cumulative effect of several exclusion rules stacked together is to trim the tails of the reported distribution somewhat, a detail that rarely makes it into a summary of a study's headline mean.
Who does the measuring
This is the single decision Veale's inclusion rule treats as non-negotiable: only studies in which a health professional took the measurement were pooled, and self-reported figures were excluded regardless of sample size. The reasoning is direct - self-reported figures run consistently and substantially higher than clinician-measured ones, for reasons that are more about method and rounding than deliberate exaggeration, and pooling the two together would produce a figure that represents neither population cleanly. A study's design decision on this single point is often the largest single factor separating its reported mean from another study's, larger than genuine population differences.
What instrument and method were used
Even among clinician-measured studies, instrument and method vary. Most used ordinary rigid rulers for length and flexible tapes for circumference - nothing exotic - but whether length was taken bone-pressed or not, and whether it was measured along the dorsal line specifically, differs between older and newer studies and is not always stated as clearly as it should be. A change in protocol between bone-pressed and non-bone-pressed method, or between stretched and erect state, can shift an entire study's reported mean by an amount that swamps its confidence interval - which is why the method line matters as much for a published study as it does for a reading you take yourself.
What state participants were measured in
Flaccid, stretched flaccid, and erect are three different measurements with three different means and three different amounts of spread, and a study's design has to specify which one it obtained and how. How studies actually obtain a genuine erect measurement matters here specifically, because the options - pharmacologically induced erection in a clinical setting, private self-stimulation, or self-report of a home reading taken under instruction - are not interchangeable, and each produces a somewhat different dataset even when every other design choice is held constant. This is also the reason erect samples in the pooled literature are consistently smaller than flaccid ones: fewer studies were able to obtain a genuine erect measurement under controlled conditions at all.
How the result gets reported
The final decision is reporting: does the paper give a mean and standard deviation, a mean and confidence interval, or a bare average with no spread at all? These are different statements and get confused constantly - an SD describes the individuals in the sample, a CI bounds where the true mean probably sits, and a paper that reports only one of them is telling you less than a paper that reports both. A study whose design was otherwise sound but whose reporting drops the spread is far less useful for anything beyond citing the single mean, because a mean without a spread cannot support a percentile, an uncertainty statement, or a meaningful comparison to another study.
What repeated measurement within the study buys
A well-designed study does not stop at one reading per participant, and how many it takes, and by whom, is itself a design decision worth checking. Studies with two independent observers measuring the same men report agreement that is good but not perfect, which sets a practical floor under how precise even the best clinician-measured figure can be - if two trained professionals measuring the same person do not land on exactly the same number, no single reading from any protocol should be treated as exact. That floor matters when a study's reported precision is compared to its actual design: a mean reported to one decimal place is not claiming millimetre-level truth about any individual, it is a statement about central tendency across a sample, and rounding conventions in published tables are themselves a signal about how much precision the data can honestly support.
What a small study's design costs it
Sample size interacts with all five of the decisions above rather than sitting apart from them. A study of a few dozen men can report a mean a full centimetre away from a larger pooled estimate purely through sampling variation, with no error in its method at all - smaller samples are simply noisier, and their own confidence intervals say so plainly if a reader checks them rather than taking the headline mean at face value. This is one more reason a meta-analysis that pools many studies, weighting them by size and quality, produces a more stable reference than any single smaller study can, however careful that single study's individual protocol was.
What era a study was run in
Design choices are not fixed across time, and a study run decades ago was operating under different norms for several of the six decisions above than one run recently. Kinsey's mid-century figures, gathered by postcard and self-measured entirely, are the founding example of a design that would not meet a modern inclusion bar - no clinician measurer, no controlled state, no verification of any kind - and yet a rounded version of that era's numbers is still what a lot of casual conversation quietly assumes as a baseline. A single small study from 1996 established the case for using stretched flaccid length as a practical proxy for erect length by measuring the same men flaccid, stretched and erect and comparing the three figures directly, a design choice - measuring one population three ways rather than three separate populations once each - that is exactly why the stretched-to-erect relationship in the modern pooled data can be trusted as a within-person comparison rather than an across-study coincidence. Knowing roughly when a study was run, alongside its other five design choices, is a fast way to guess at which norms it was likely following before checking its methods section directly.
Reading a study through this lens
A well-designed study answers a length question about a defined population, and it does not do more than that no matter how carefully every one of the six decisions above was made. Put the six decisions together and a study's reported figure stops looking like a fact and starts looking like the output of a specific, inspectable process: recruit from where, exclude whom, measured by whom, with what instrument, in what state, reported how. Two studies that disagree by a centimetre are not necessarily measuring different populations - they may simply have made different choices at two or three of these six points, and the disagreement is fully explained before anatomy enters the picture at all. This is also the correct way to read Veale itself: its own explicit choice on the measurer question is exactly why its figures read lower than casually sourced averages, and its smaller erect sample is exactly a consequence of how hard the state decision is to satisfy well.
None of these design questions have anything to do with how a photograph gets scored. A rating from Rate Cock, the model behind AI Penis, the aggregate scores Penis Rater publishes, or a person's judgement at Rate Penis is built on a rubric and a rater, not on a recruitment protocol or a measuring instrument, and none of the six decisions above apply to it in any form. Knowing how a measurement study is built is what lets you read one skeptically without becoming skeptical of measurement itself - the process above is exactly what makes a figure like Veale's trustworthy in the first place, once you know which choices it made and why. That is a different kind of trust from the kind a rating earns, which rests on a rubric being applied consistently rather than on a recruitment protocol being sound - two questions worth keeping separate whenever a number from either world gets quoted as though it settled something the other could not.