Data

Reading a nomogram without fooling yourself

What a percentile is, how a nomogram is built from a mean and a standard deviation, and how to place a measured figure on it with its uncertainty attached.

7 min readData

A percentile is one of the most misread numbers on this whole subject, not because the arithmetic is hard but because a single figure gets treated as more precise than the process that produced it can support. This is the full mechanics: what the number means, how it is built from published data, and how to place your own reading on it honestly.

What a percentile actually is

A percentile states the share of a reference sample that measured below a given figure. The 50th percentile is the median: half the sample measured less, half measured more. The 90th percentile means nine in ten men in that reference sample measured a smaller figure than the one being placed - it says nothing about the other one in ten beyond "larger," and nothing at all about anyone outside the reference sample. What each specific percentile value means in plain terms, worked through the common ones, is covered on its own; the point that matters here is that a percentile is always relative to a stated population, never an absolute property of a number.

Where the underlying curve comes from

Veale et al. (2015) reports a mean and a standard deviation for each measure it pools - flaccid length, stretched length, erect length, flaccid circumference, erect circumference. A nomogram assumes the underlying population is approximately normally distributed around that mean, with a spread described by the standard deviation, and builds a curve from those two numbers alone. This is a real assumption, not a formality: the authors checked it against their pooled data and found it held reasonably well, though less exactly at the far tails, which is worth understanding on its own before trusting the extreme ends of any chart built this way.

Why a percentile needs a named reference population

A percentile is meaningless without saying which sample it is measured against, and this is the single most common thing a casual percentile claim leaves out. The 90th percentile against Veale's clinician-measured pool is a different figure entirely from the 90th percentile against a self-selected online survey, because the two underlying distributions have different means and different spreads - a self-report sample runs higher on average for the reasons covered in full elsewhere, so the same physical length lands at a lower percentile against it than against a clinician-measured reference. Any percentile worth citing should name its source study in the same sentence, the way this site tries to do consistently, rather than presenting the number as though percentiles were a single universal scale. The same reference-population caveat applies across states too - how flaccid and erect percentiles compare is worth reading before assuming a percentile computed for one state says anything about the other.

A related, narrower term worth being precise about is "normal range." In a results table, that phrase is a statistical description - typically the middle 95 percent of a sample, two standard deviations either side of the mean - not a judgement about anyone outside it. A figure outside that band is unusual relative to the sample, which is a factual statement about frequency, and carries none of the evaluative weight the everyday use of the word "normal" tends to imply.

From a measurement to a z-score to a percentile

The actual arithmetic is three steps, and every nomogram is doing this same calculation whether or not it shows its working.

Subtract the mean. Take your figure and subtract the reference mean for that measure and state - for erect length against Veale, that mean is 13.12 cm.

Divide by the standard deviation. The result is a z-score: how many standard deviations your figure sits from the mean. For erect length, the SD is 1.66 cm, so a figure of 14.78 cm sits exactly one SD above the mean, a z-score of 1.

Look up the z-score against the normal distribution. A z-score of 1 corresponds to roughly the 84th percentile under a normal curve; a z-score of 2 corresponds to roughly the 98th; a negative z-score simply means the figure sits below the mean, and the same lookup applies with the sign flipped. None of these percentages are specific to this subject - they come from the shape of the normal curve itself, and the same lookup table works whether the underlying measure is a length, a height, or a test score, as long as the underlying distribution is genuinely close to normal. This mapping, worked through with the actual numbers rather than described abstractly, is covered step by step separately, because seeing it done once with real figures is worth more than the formula alone.

A nomogram is this same lookup drawn as a curve, so a reader can go from a measured figure straight to an approximate percentile without doing the division by hand.

Why the middle of the curve is crowded

Under a normal distribution, values pile up near the mean and thin out toward the tails. A relatively small error near the mean therefore shifts a percentile by far more than the same error would near an extreme - a centimetre of measurement error close to average length can move the estimated percentile by ten points or more, while the same centimetre of error at the far tail barely moves the percentile at all, because there is so little of the distribution left to move through. This is a direct, unavoidable consequence of the shape of the curve, and it means a percentile derived from a sloppy reading near the middle of the distribution is far less trustworthy than the same sloppy reading would be at an extreme.

It also explains why the tails of the distribution are thinner than casual talk about them implies: every additional standard deviation out cuts the remaining share of the population sharply, so claims about rare, far-tail figures should be read with proportionally more scepticism than claims near the middle.

Placing your own figure on the curve, honestly

The step most people skip is carrying an uncertainty through the calculation rather than treating a single reading as exact. If careful technique gets you to something like ±0.3 cm on a length reading, that uncertainty translates into a range of z-scores, not a single one, and therefore a range of percentiles rather than one precise-looking figure. A reading of 13.4 cm ± 0.3 cm against the erect mean and SD above works out to a percentile band, not a point - somewhere in the high 50s to mid 60s, roughly, rather than a single number that implies more precision than three readings on a home ruler can actually support. Reporting the band, or at minimum acknowledging it exists, is the difference between using a nomogram honestly and using it to manufacture false confidence in a figure that was never that precise to begin with. The same logic applies twice over if the figure being placed came from a single reading rather than a median of several - a single reading carries the full technique uncertainty on top of the instrument's own error bar, and a percentile built from it should be treated as looser still than one built from a properly repeated measurement.

The rule that governs the whole curve

The normal distribution that a nomogram assumes follows a fixed pattern regardless of which measure it is applied to: about two-thirds of a sample falls within one standard deviation of the mean, and about 95 percent falls within two. Applied to erect length against Veale's figures, that means roughly two-thirds of measured men fall between about 11.5 cm and 14.8 cm, and about 95 percent fall between about 9.8 cm and 16.4 cm - a concrete range built directly from the mean and SD in the table, rather than an abstract statistical rule. At the far end of that curve, the 99th percentile sits about 2.33 standard deviations above the mean under the normal assumption, which for erect length against Veale works out to a specific, checkable figure - and it is worth doing that arithmetic yourself before accepting an extreme claim at face value, because very few real figures survive contact with what the 99th percentile actually requires.

Why calculators built from this disagree with each other

If you have ever run the same measurement through two different online percentile calculators and gotten two different answers, the reasons are almost always in the inputs, not the arithmetic: a different reference study, a different assumed state, or a tool that silently accepts a self-reported figure as though it were clinician-measured. The specific reasons two calculators can diverge, and what a trustworthy one discloses about its own assumptions, are covered in full elsewhere and what a calculator worth using actually states up front - the short version is: check what reference population and state a tool assumes before trusting the number it returns.

What a percentile is not

A percentile answers "where does this length sit against a reference population of measured men," and nothing else. It is not a rating, a score, or a prediction about how a photograph would be judged - that is a separate kind of number built from a rubric and a judge rather than a mean and a standard deviation, the kind Rate Cock works with, or that a model at AI Penis or the published aggregate scores at Penis Rater produce from an image, or that a person gives directly at Rate Penis. A percentile from a nomogram like this one is a statement about a distance next to other distances, computed from a named study with a stated mean and spread, and its honesty depends entirely on carrying that study's own uncertainty along with the figure rather than dropping it at the first decimal place.

Read next

Full archive