Data
The world map of averages
The country-by-country averages that circulate online mix self-reported surveys with clinical studies of different methods, then rank them to one decimal place.
A map or table ranking countries by average length circulates periodically, presented with a confidence that nothing about its underlying data supports. The problem is not that a single number is wrong - it is that the table's rows are not comparable to each other, and the ranking treats them as if they were.
How these tables actually get assembled
A country league table needs one figure per country, and the sources available for any given country are not uniform. Some countries have a clinician-measured study behind their entry; many do not, and the entry is filled from a self-reported survey instead, or from an older figure of unclear provenance that has simply been repeated by enough secondary sources to look authoritative. Even among the entries that do trace back to a measured study, those studies differ in state (flaccid, stretched, erect), in whether the ruler was bone-pressed, in sample (clinic-recruited versus general population), and in sample size, which for some countries is small enough that the reported mean carries a wide margin nobody displays alongside it. The table then places all of these, regardless of origin, into a single ranked list, usually to one decimal place, which implies a level of precision and comparability that no single figure in the list actually earned.
Why mixing methods breaks a ranking specifically
A difference of a centimetre or more between two studies can come entirely from the reasons two clinician-measured studies disagree in general - state, bone-pressed status, posture, sample - before any real difference between countries is involved. Add self-reported entries into the same table, and the size of that problem grows further, since self-reported figures run consistently higher than measured ones for reasons that have nothing to do with geography. A country whose entry happens to come from a self-reported survey will rank higher than an otherwise-identical country whose entry comes from a clinician-measured study, purely because of which kind of source was available, not because of any real difference between the two populations. A ranking is especially sensitive to this kind of noise, because it converts small, method-driven differences in the underlying figures into an ordered list that reads as a strong claim even when the gaps between adjacent rows are well within the range that protocol differences alone can produce.
What a defensible comparison would need
A country comparison worth trusting would need the same measurement protocol applied by trained observers across every country included, similar recruitment methods so the samples are comparable, sample sizes large enough that the resulting margins are usefully tight, and a reported uncertainty next to every figure rather than a bare mean. No table circulating informally meets that bar, and building one that does would be a substantial piece of original clinical research, not a spreadsheet exercise. Absent that, the responsible reading of any country league table is that it shows which countries happen to have which kind of data available, more than it shows a real ranking of anatomy.
This is a related but distinct problem from what Veale 2015 could and could not say about ethnicity - that paper's limitation was insufficient data within a single pooled, clinician-measured meta-analysis; a country league table's problem is actively mixing sources of different quality and calling the result a ranking.
None of this resembles a subjective assessment, and no country table should be mistaken for one. A rating platform, an image-based estimate, a human judge's score, or a gallery ranked by score is not drawing on population data by country at all, and treating any of them as evidence for or against a country's place on a league table would be a category error in the other direction.