Data
The Veale meta-analysis, read properly
The paper the whole subject rests on: what it set out to do, what it pooled, what it reports, and where its own authors say to be careful.
Almost every figure repeated on this subject, correctly or not, traces back to one paper. Veale, Miles, Bramley, Muir and Hodsoll, published in BJU International in 2015 under the title "Am I normal? A systematic review and construction of nomograms for flaccid and erect penis length and circumference in up to 15,521 men", is a systematic review and meta-analysis rather than a single study, and reading it properly means understanding what a meta-analysis of this kind can and cannot do before looking at a single number it reports.
What the paper set out to do
The authors were responding to a specific gap: a lot of men who believe themselves to be below average have no reliable reference to check that belief against, because the figures circulating publicly are a mix of self-report surveys, single small clinical studies, and numbers with no traceable source at all. The paper's stated aim was to pool every study it could find that met a defined quality bar, combine them into single reference figures, and build nomograms - charts that let a measured figure be placed against a percentile - for both flaccid and erect length, and for circumference.
The inclusion rule that shapes everything else
The single decision that matters most in reading this paper is its inclusion criterion: only studies where the measurement was taken by a health professional, using a standard method, were pooled. Self-reported figures, however large the sample, were excluded outright. This is the reason the paper's figures read lower than the numbers people are used to seeing quoted online - not because the authors chose a flattering study, but because they systematically removed the self-report inflation that dominates the loosely sourced averages in circulation. How that inclusion rule was applied, and what it left out as a consequence, gets its own paragraph below and its own post.
Twenty studies met the criteria, contributing a combined sample of up to 15,521 men, though not every man in that total was measured on every outcome - flaccid length and flaccid circumference had the largest pooled samples, erect measurements the smallest, because far fewer studies measured men in an erect state under clinical conditions.
How the review actually found its studies
A systematic review lives or dies on its search strategy, and this one followed the standard approach for the field: a structured search of the medical literature databases, screening every result against the inclusion criteria above, and a manual check of the reference lists of studies that passed, to catch anything the database search missed. Studies that only reported self-measured or self-reported figures were screened out at this stage regardless of how large or well-known they were, which is why some frequently cited older surveys do not appear anywhere in the pooled figures - they simply did not meet the bar the review set for itself before looking at a single number. This screening step is invisible in the final table, but it is where most of the paper's credibility actually comes from: a table of means means little without knowing what was excluded to produce it and on what grounds.
Why this paper displaced the studies that came before it
Before 2015, anyone wanting a reference figure had to reach for a single study - smaller, older, and each carrying its own specific population and method - and different single studies disagreed with each other by amounts that were hard to explain without a broader view of the field. A meta-analysis does not remove that disagreement between the underlying studies; it makes the disagreement visible and quantifiable rather than hidden inside a single number, which is a genuine improvement even though it means the pooled figure carries a note of caution that a single confident-sounding study did not. What a pool of twenty studies buys, specifically, is precision about the mean - a large combined sample narrows the confidence interval around the average considerably - without buying anything at all about the spread of individuals, which is set by the standard deviation regardless of how many people were measured to establish it. That distinction, between what a big sample does and does not tighten, is easy to state and easy to miss in practice, and it is worth holding onto separately from the headline sample size whenever a figure this large gets quoted as though it settled the matter completely.
The pooled figures
The headline numbers, as the paper reports them, are means with standard deviations attached rather than bare averages:
| Measure | Mean | SD | Approx. n |
|---|---|---|---|
| Flaccid length | 9.16 cm | 1.57 cm | 10,704 |
| Flaccid circumference | 9.31 cm | 0.90 cm | 9,407 |
| Stretched flaccid length | 13.24 cm | 1.89 cm | 3,821 |
| Erect length | 13.12 cm | 1.66 cm | 692 |
| Erect circumference | 11.66 cm | 1.10 cm | 381 |
A few things are worth reading directly off that table rather than off a summary of it. Stretched flaccid length and erect length sit close together, which is the empirical basis for using a stretched reading as a practical stand-in for an erect one in settings where obtaining a genuine erect measurement is impractical. Circumference carries a noticeably tighter spread, relative to its mean, than length does in either state - a pattern that shows up again when the two are compared as percentiles, where the same absolute error in a girth reading moves a percentile further than the same error in a length reading, because the population is more tightly clustered on that measure.
What "explained in full" needs to include: the nomograms
A mean and an SD are the raw material; the paper's actual deliverable, the thing that made it useful beyond a table, is the set of nomograms it constructed from them - charts that let a reader take a measured figure and read off roughly where it sits against the pooled sample, assuming the sample is approximately normally distributed. Each nomogram, and what its curve actually plots, is covered in full separately, because the construction method deserves its own explanation rather than a compressed paragraph here. The short version: a nomogram built this way is only as good as the assumption that the underlying distribution is close to normal, which the authors checked and reported was reasonable for their pooled data, though not exact at the extreme tails.
Where the authors themselves say to be careful
A meta-analysis inherits the limitations of the studies it pools, and this paper's discussion section is explicit about several of them rather than glossing past them. The pooled studies were heterogeneous in population - clinic-based samples in several cases, rather than a genuine general-population draw - which the authors flag as a reason the mean may not transfer cleanly to every population a reader belongs to. Erect data came from far fewer men than flaccid data, which widens the practical uncertainty on the erect figures even though the table above reports them with the same apparent precision as the rest. And the pooled studies did not report ethnicity consistently enough for the authors to draw any conclusion about differences by ethnicity, a gap they name directly rather than filling with an assumption. The full list of stated limitations, in the authors' own terms, is worth reading on its own rather than compressed into a caveat here, because several of them change how much weight a specific figure in the table above should carry.
There is also a structural limitation that is easy to overlook because it is built into what a meta-analysis is rather than something the authors could have designed around: pooling twenty studies means pooling twenty slightly different populations, protocols and time periods into a single mean, and when those studies disagree by more than sampling variation alone would predict, the pooled figure is an average of genuinely different things rather than one thing measured twenty times. Veale and colleagues report a formal measure of this heterogeneity alongside their pooled means, which is exactly the kind of detail a headline figure strips away and a careful reading restores.
Reading the sample sizes as information, not decoration
The "approx. n" column in the table above is not incidental. A pooled mean built from 10,704 flaccid measurements carries a narrower confidence interval than one built from 381 erect circumference measurements, even though both are reported to the same number of decimal places in the summary table. Erect data specifically are thinner than flaccid data throughout this literature, for the practical reason that obtaining a genuine erect measurement under clinical conditions is harder to arrange than a flaccid one, and that gap in sample size is a standing reason to treat the erect figures as somewhat less pinned down than their flaccid counterparts, table formatting notwithstanding.
Girth gets its own note
Circumference is reported more sparsely than length throughout the pooled literature, for reasons that are mostly about which studies bothered to measure it and under what state, and the erect circumference sample above is markedly smaller than any of the length samples. What the paper reports for girth specifically, and what that smaller sample implies about how much to trust the figure, is worth its own read rather than a line in a summary table.
Reading the paper rather than a repost of it
The single most common failure mode with this paper is citing a figure from it secondhand, via a chart or an app that has redrawn the nomogram without saying so, rather than from the paper or a source that traces back to it cleanly. How to cite it correctly, and which figure belongs to which outcome, is covered separately, and is worth doing properly if a number from this paper is going into anything you intend to stand behind.
None of what this paper measures maps onto a subjective score. Percentiles built from these figures answer a length question, not a judgement question - which is a distinction worth keeping in mind before treating a number from this table as though it predicted how a photo might be rated on a site like Rate Cock, scored by the model behind AI Penis, aggregated the way Penis Rater does, or judged directly by a person at Rate Penis. Those are different instruments answering a different question, and this paper does not speak to any of them.