Data
Heterogeneity
When the studies in a pool differ more than chance allows, the pooled mean is an average of different things; Veale 2015 reports this and it should be read.
Pool enough studies and they will not agree perfectly, even if every one of them measured the same true quantity with no bias at all. Some disagreement is expected from sampling alone. Heterogeneity is the term for disagreement beyond what sampling alone explains, and it is a specific, checkable property of a meta-analysis rather than a vague complaint about the data.
What heterogeneity means, conceptually
Imagine five studies each measuring the same fixed quantity in a large enough sample that pure chance would produce results clustered tightly around the true value. If the five results instead scatter more widely than that chance-based clustering predicts, something beyond sampling noise is driving the difference between them - different populations, different protocols, different measurers, or some mix of all three. That extra scatter is heterogeneity, and meta-analyses report a statistic for it precisely so a reader does not have to guess whether it is present.
Why it arises here
A pooled estimate like Veale et al. (2015) draws on studies conducted in different countries, different clinical settings, and across a span of years, using measurement protocols that were broadly similar but not identical in every detail - bone-pressed versus non-bone-pressed length, for instance, is a real methodological split across the older literature, not a uniform standard from the start. The recruited populations differ too: some studies drew from urology clinics, others from different clinical contexts, which is its own source of variation before method is even considered. Given that, some degree of heterogeneity in the pooled figures is close to expected rather than surprising, and the paper's authors report it rather than hiding it.
What it changes about reading a pooled mean
When heterogeneity is present, a pooled mean is best understood as an average across somewhat different measurement exercises rather than as many repeated readings of one identical thing. That does not make the mean useless - it is still the best available central estimate, drawn from the largest sample anyone has assembled on this question - but it changes the precision you should credit it with. A confidence interval around a pooled mean, in the presence of real heterogeneity, understates the uncertainty if it is read as though every constituent study measured under identical conditions. What a large pooled sample size does and does not buy you is a related but distinct point: sample size tightens the estimate of the mean itself, while heterogeneity is a separate question about how comparable the things being averaged actually were.
Heterogeneity is not the same as bias
It is worth being precise about what heterogeneity is not. It is not evidence that any individual study was measured badly, and it is not the same claim as self-report inflation, which is a distinct and much larger effect this site covers on its own. It is simply a statement that the studies in the pool do not agree as tightly as sampling alone would predict, for reasons that can include method, population, or measurer, without implying that any one of them is wrong.
Reading Veale 2015 with this in mind
The paper names heterogeneity among its own stated limitations, alongside the other caveats worth carrying with the headline figures - the full list of what the authors flag is worth reading alongside the means themselves. The practical upshot for a reader comparing their own measurement to the published figure: treat the pooled mean as a solid central estimate rather than a precise physical constant, and hold the confidence interval around it a little more loosely than the paper's own arithmetic would suggest in isolation.
None of this bears on a photograph-based judgement of any kind - a rating from Rate Cock is not a pooled statistical estimate and does not carry a heterogeneity statistic, because it is assessing one image rather than averaging measurements across studies. The same goes for an automated read from an image model, which is not doing anything resembling meta-analysis when it scores a photograph, or a score from Penis Rater, which ranks images against each other rather than pooling clinical data. A verdict from a human judge on Rate Penis is a single opinion, not an average of several measurements at all, which is a different kind of statement than anything a meta-analysis produces.