Data

A large sample tightens the mean, not the spread

A big sample pins down the average precisely; it does nothing to narrow the range of individuals, which is set by the standard deviation.

By 3 min readData

Guides on Data: How a size study is built, Reading a nomogram without fooling yourself, The Veale meta-analysis, read properly

Veale et al. (2015) pools measurements from 15,521 men, and the size of that number gets treated as a general guarantee of accuracy. It is a guarantee of one specific thing, and not of the thing most readers assume.

What a large sample actually improves

A bigger sample narrows the standard error of the mean - the uncertainty around where the true population average sits. With 15,521 men, the pooled mean length is known to a precision that a study of a few dozen men could never approach; if you repeated this exercise on a different but comparably drawn sample of men, the resulting mean would land very close to the one already published. That is what a large sample buys: confidence in the location of the average.

What it does not touch

It says nothing about how spread out individual men are around that average. That spread is described by the standard deviation, and the standard deviation of a population does not shrink as you measure more people - it is a property of the population itself, not of how well you have estimated its mean. Measuring 15,521 men instead of 500 tells you the average more precisely; it does not make individual men cluster more tightly around it, because they were never going to.

This is a common confusion and worth stating plainly: a huge sample size is not evidence that everyone is close to average. Whether people are close to average or widely spread is a separate fact, reported by the standard deviation, and a sample size in the thousands has no bearing on it either way.

Why this matters for reading your own figure

If you are placing a personal measurement against the published data, the number that actually tells you something about where you sit is the standard deviation, not the sample size. A large N tells you the mean and the SD are both estimated reliably - you can trust that the reported spread is close to the population's real spread, because 15,521 men is enough to pin down a standard deviation as well as a mean. But it is the SD itself, not the size of the sample behind it, that determines how far from average any given measurement actually is.

A concrete way to hold both facts

Think of it as two separate questions the paper answers with different tools. "Where is the average?" is answered precisely because of the large sample - the confidence interval around the mean is narrow. "How different are individual men from each other?" is answered by the standard deviation, and that answer would be the same whether the paper had pooled 500 men or 15,521, give or take some noise in estimating it, because it describes the underlying population rather than the precision of the estimate. Confusing the two is why a headline sample size gets used to imply something about individual variation that it was never measuring.

Where to go from here

For the mechanics of that second question - what the standard deviation actually represents and how to read it against a single figure - see standard deviation, explained with length, and for the underlying figures this whole discussion is built on, what the published data actually says reports the full table.

None of this scales up into a statement about any one photograph, either. A rating from Rate Cock is an assessment of a single image, not a pooled statistical estimate, so questions of sample size and standard deviation do not transfer across to it. An image model has the same limitation in reverse - it cannot pool anything from a single photograph, because there is no dataset behind the number it produces the way there is behind Veale 2015. A scored gallery like Penis Rater ranks images against each other rather than estimating a population mean, and a human judge's opinion is a single data point, not a sample at all.

The confusion between sample size and individual spread is common enough to be worth restating plainly, because it recurs anywhere a headline number gets quoted without its second statistic attached. A large N is a statement of confidence about a location on a chart, not a statement about how tightly everyone clusters around it - and a reader who takes the first to mean the second will misjudge how unusual their own figure actually is, in either direction.

Read next

Full archive