Data

What the authors say to be careful about

The paper's own limitations section names heterogeneity, clinic-based samples, uneven erect data and unverifiable ethnicity; a reader should carry them along with the means.

3 min readData

Veale et al. (2015) is the standard reference this site cites throughout, and it earns that status partly by being unusually clear about its own limits. A reader who quotes the means without the caveats the authors themselves list is quoting the paper less carefully than the authors wrote it.

Heterogeneity between the pooled studies

The seventeen studies pooled into the meta-analysis did not use identical protocols, populations or measurers, and the paper reports that the studies differ from each other by more than sampling chance alone would explain. That matters because a pooled mean assumes, to some degree, that it is averaging comparable things. When heterogeneity is present, the pooled figure is closer to an average of somewhat different measurements than a single precise estimate of one quantity - what that actually does to a pooled mean is worth understanding on its own, separately from this walk through the paper's stated limitations.

The samples are clinic-based

Most of the men in the pooled studies were recruited through urology or sexual-health clinics, not drawn at random from the general population. Men attending a clinic for an unrelated reason are not guaranteed to resemble the population at large, and the authors note this as a limit on how far the figures generalise. Clinic samples and the general population are not the same thing, and the direction of any resulting bias is not something the paper claims to know either.

Erect data are the thinnest slice

Far fewer men across the pooled studies were measured in a fully erect state than were measured flaccid or stretched flaccid. Erect measurement is harder to standardise in a clinical setting, so protocols leaned on the flaccid and stretched states more heavily, and the erect figures accordingly carry wider uncertainty than the other two states, a gap the authors acknowledge directly.

Ethnicity could not be assessed

The paper looked for ethnic differences in the pooled data and states plainly that the data were insufficient to draw a conclusion either way. That has not stopped the paper being cited in both directions on the question since publication, which is exactly the kind of misuse a stated limitation is supposed to prevent. What the meta-analysis could and could not say about ethnicity is worth reading in full if that is the question you came with, because the honest answer is an absence of evidence rather than a finding.

Why this list matters more than the table

A mean without its limitations is a number stripped of the conditions that make it interpretable. The four items above are not reasons to discard the paper - it remains the best available reference for this reason: it is the largest, most methodologically consistent pool that exists, measured by professionals rather than self-reported. They are reasons to hold the specific figures with the caveats the people who produced them attached, rather than repeating a mean to two decimal places as though it arrived without conditions.

None of these limitations have anything to do with a subjective assessment of a photograph, which is a different kind of statement altogether - the sort Rate Cock produces from an image rather than a tape. An automated version of that same judgement runs into its own separate limit regardless of what the clinical literature says - an image model has no access to a physical scale in a photograph, so none of Veale's caveats about sampling or heterogeneity are even the relevant kind of limitation for that comparison. A scored gallery like Penis Rater and a human opinion from a judge on Rate Penis both sit outside this paper's scope entirely, which is worth saying plainly rather than leaving implied.

For the full walkthrough of what the paper set out to do and what it pooled, this post is deliberately narrower than a complete summary; it exists to carry the four caveats above alongside the figures wherever they get quoted.

Carrying them costs nothing. Attaching "clinic-based, uneven erect data, unresolved ethnicity, some heterogeneity between studies" to a quoted mean takes one sentence, and it is the difference between citing the paper and citing a number that has been quietly detached from the conditions under which it was produced. The authors did the work of naming these limits themselves; the least a reader can do is keep them attached.

Read next

Full archive