Data

How Kinsey collected size data

Kinsey's figures were self-measured by participants and mailed back on cards, which makes them the founding example of the self-report problem.

4 min readData

Alfred Kinsey's mid-twentieth-century surveys of male sexual behaviour are the earliest large dataset on this subject that most people have heard of, and the method behind them explains their figures better than anything about the men involved. Participants measured themselves at home and reported the result, in some cases on a card returned by post. No clinician verified the reading, checked the technique, or observed the measurement at all.

Why that was a reasonable design at the time

It is easy to judge a seventy-year-old survey method against a standard it was never trying to meet. There was no clinical protocol for this measurement in wide use at the time, no large-scale infrastructure for observed measurement at the scale Kinsey was working, and mailed self-report was a genuinely practical way to collect data from thousands of participants who would never have agreed to a clinic visit for the purpose. For a study of sexual behaviour broadly, rather than a study built specifically to measure size, self-report was a defensible trade-off between data volume and data precision.

Where the method breaks down for this specific figure

Self-measurement without a checked protocol fails in ways that are well understood now. A participant with no instruction to press to the pubic bone will not press to the pubic bone - that single step is what removes most of the difference between a self-measurement and a clinical one, and nobody mailing in a card had been told to do it. Nobody checked which line was measured, top or underside. Nobody confirmed the state of erection was full and constant, and nobody vetted the ruler or the reading against a second observer.

Each of these is a source of upward bias individually, and they compound. The same mechanisms - method drift, upward rounding, remembering the best of several attempts - are documented in modern self-report surveys, and there is no reason to think a mid-century postcard survey was exempt from any of them. If anything, with zero instruction on technique, it is a cleaner example of the problem than most.

Why it is not comparable to the modern clinical pool

Veale et al. (2015), BJU International, pooled 17 studies and set its inclusion bar at measurement by a health professional using a stated method. Kinsey-era data fails that bar entirely, not because the participants were dishonest, but because nothing about the collection method could produce a figure comparable to a bone-pressed clinical reading. Comparing the two is comparing different measurements that happen to share a unit.

This is not a criticism unique to Kinsey. It is the general reason the folk figures still in circulation trace to self-report rather than to clinical data, and Kinsey's postcard method is simply the earliest large-scale version of the pattern that later self-report surveys, online and offline, have repeated ever since. That mailed-card pattern is structurally identical to what happens on an anonymous online survey today, just without the internet's speed or reach.

A method problem, not a character problem

None of this requires assuming participants exaggerated on purpose. The mechanisms that inflate a self-measured figure operate whether or not anyone intends to mislead: a ruler not pressed to bone reads long by itself, an ambiguous instruction gets interpreted generously by default, and a number written down under no supervision at all has no check on it either way. Treating the gap between Kinsey-era figures and later clinical ones as evidence of dishonesty misreads what actually happened, and it is a more forgiving, more accurate reading of the data to treat it as a method finding instead.

What the postcard method is actually good evidence of

Read for what it is, this data is a genuinely useful early record of what men reported about themselves without correction - which turns out to be a more durable and interesting finding than the raw figures themselves. It is a data point about self-report, not a data point about anatomy.

That distinction matters more broadly than statistics. A self-reported figure, an AI-estimated figure, and a subjective rating are three different kinds of number that only look alike because they share units or a scale. An image model reading a photograph is inferring proportion from visual cues, not recovering a measurement, which is closer in kind to a mailed-in guess than to a bone-pressed reading. A rating out of ten from a service like Rate Cock is a different kind of number again - a judgement, not a distance - and a human judge's verdict sits in that same category rather than in the one Kinsey's participants were trying, and failing, to occupy. A public distribution of scores on a rating hub is closer to the postcard method than it looks at first: both are collecting a great many self-selected submissions, and neither one is a clinical measurement however large the pile gets.

Read next

Full archive