Data

Designing the honest comparison

The clean test of self-report inflation is the same men measured both ways under one protocol; here is what that design needs and what has come close.

3 min readData

The claim that self-reported size runs high rests on comparing two different groups of men measured two different ways. That is suggestive, not proof. The clean test looks different, and almost nobody has run it in full.

What the comparison actually needs

A proper test of the self-report gap needs the same individual measured twice: once by himself, at home, the way a survey respondent would, and once by a trained observer using the clinical protocol, taken the way the literature does it. Both readings need to come from the same man, close enough in time that his anatomy has not changed, far enough apart that one reading does not anchor the other.

Order matters more than it looks like it should. If the self-measurement happens first, the participant carries that number into the clinical visit and may unconsciously position himself, or the observer may unconsciously round toward it. If the clinical measurement happens first, the same problem runs the other way. A design that cares about the gap randomises which one comes first, and does not tell the observer what the participant reported.

Blinding is the part almost every comparison skips

Most of what circulates as evidence for self-report inflation is a mean from a self-report survey sitting next to a mean from a clinician-measured study, cited as if they describe the same population. They do not. Different men volunteered for each, under different incentives, and nobody controlled for anatomy, only for the population average. A cross-study comparison like that tells you the two literatures disagree. It cannot tell you why, because population, motive and method are all confounded together.

The design that actually isolates self-report as a cause is a within-subject one: one man, two readings, one clinical and one self-taken, with the observer blind to the self-report and, ideally, the participant not told his clinical figure until after he has reported his own. Without blinding, either number can drift toward the other and the gap you measure is smaller than the real one.

What has come close

A handful of studies have paired a self-reported or self-taken figure against a clinician's reading in the same visit, generally as a secondary check on a study's main measurement rather than as the primary question. Where this has been done, the self-taken figure tends to run higher than the clinician's, which is consistent with the direction reported elsewhere, but the sample sizes in these paired comparisons are small next to the survey literature, and few of them blinded the observer to the self-report in the way a clean test would require. That leaves the size of the gap less settled than the direction of it. Veale et al. (2015), BJU International, the standard reference for the clinician-measured pool, did not set out to test self-report against clinical measurement in paired individuals - it pooled clinician-measured studies precisely to exclude self-report from the estimate, which is a different and narrower goal than the paired design described here.

Why the substitute is still worth reading

None of this means the population-level comparison is worthless. A self-report survey and a clinician-measured meta-analysis disagreeing by a consistent margin, across many independent samples, is itself evidence, just weaker evidence than a paired design would give. It rules out the possibility that the gap is a one-off sampling accident, even if it cannot pin down the exact contribution of any single mechanism - the mechanisms behind self-reported figures running high are covered on their own. Reading a table of study means, from either kind of study, means asking what was actually compared, not just what the numbers say - the error column next to a mean carries information the mean alone does not.

A photograph carries no scale information at all, so an AI model estimating from an image is not even in the self-report-versus-clinical comparison - what an image model is actually inferring instead is a third kind of number with its own error source entirely, and a human judge's impression, however experienced, is a different kind of verdict again, not a third measurement to average in. None of this is a subject Rate Cock tackles either, since a photograph-based score is answering a different question than either kind of study here is asking. Where a rating service does belong in this conversation is as a reminder that a score and a length are not the same claim, and reading a score on its own terms rather than as a stand-in for a ruler is the same discipline this whole site is arguing for.

Read next

Full archive