Data
Which studies get published
Studies with striking findings are likelier to be published and cited; in this literature that favours unusual populations and surprising means.
Not every study that gets conducted gets published, and not every published study gets cited equally. Both filters push in a direction, and understanding which direction matters for reading any single figure that has survived to reach you.
What publication bias is, generally
Journals, reviewers and researchers all have some tendency to favour results that are novel, surprising, or statistically significant over results that confirm what was already assumed or find no notable effect. A study reporting an unremarkable, expected mean is less likely to attract a journal's interest than one reporting something unusual, and a null result - "we looked for a relationship and did not find one" - is historically the hardest kind of finding to get published at all. This is a well-documented pattern across many fields of research, not something specific to this subject.
Citation bias compounds it
Even among studies that do get published, some get cited far more than others. A striking or memorable figure travels through subsequent papers, popular articles and eventually forum threads, while a study reporting an ordinary result sits uncited and effectively invisible outside specialist literature searches. By the time a number reaches a general audience, it has typically passed through both filters: it survived publication, and then it survived the further selection of what gets referenced afterward.
How this plausibly acts on this subject
Applied here, both filters plausibly favour the same kind of outcome: an unusual population, an unusual method, or a surprising mean is more likely to have been published and then cited than a study that simply confirmed an unremarkable, already-expected figure. This is a statement about direction, not magnitude - there is no way to quantify from here how large the effect has been on any specific number in circulation, only to note that the mechanism exists and points one way rather than the other.
Veale et al. (2015) is less exposed to this than most sources precisely because it is a meta-analysis rather than a single striking study: it pools many studies together rather than relying on whichever individual result travelled furthest, which is part of why this site treats it as the standard reference rather than any one constituent paper in isolation. That does not make the meta-analysis immune to the bias entirely - if the pool of published studies feeding into it was itself shaped by which individual studies got published in the first place, some of that upstream selection carries through - but pooling many sources dilutes the influence of any single unusual one.
Where this shows up outside the clinical literature
The effect is much easier to see in casual, non-peer-reviewed circulation than in a paper like Veale 2015. A single online survey or self-reported dataset with a striking, high figure gets shared and requoted far more than a boring, unremarkable one, independent of which is more representative - a separate but related problem to why self-reported figures run high in the first place, which has its own, more direct causes. The country-league-table figures that circulate online show a similar pattern: sources get selected for how quotable their numbers are, not for their methodological soundness.
What this means for reading a number
The practical takeaway is not to distrust every published figure, but to weight a striking, surprising, widely-shared number more skeptically than an unremarkable one, and to prefer a pooled estimate like Veale 2015 over any single study precisely because pooling reduces exposure to this kind of selection. What a meta-analysis buys you and does not fix covers the related point that pooling does not correct every bias in the underlying studies, publication bias included.
None of this applies to a subjective rating of a single photograph, which is not a research literature and has no publication filter of the kind described here - a score from Rate Cock is an assessment of one image, not a claim competing for citation. An AI model producing a score from a photo is likewise not subject to citation bias in this sense, though it carries its own separate limitation around what it can actually read from an image that has nothing to do with publication selection. A ranked gallery like Penis Rater and a verdict from a human judge sit outside this discussion entirely, since neither is drawing on a literature that could be biased in the way a body of published studies can.