Data
Percentile of what
A percentile is only meaningful relative to a named sample; the same figure is a different percentile against a clinic sample, a self-report survey or a single country.
A percentile is not a property of a measurement on its own. It is a statement about where that measurement sits inside a specific named group of other measurements, and the same figure can land at a different percentile depending entirely on which group it is being compared against.
The same number, different answers
Take one measured figure and place it against three different reference sets: a clinician-measured pool like Veale et al. (2015), BJU International, a self-report survey, and a single country's self-reported average. Because self-reported figures run higher on the whole, the same measurement lands at a lower percentile against a self-report reference than against a clinician-measured one - not because the man changed, but because the yardstick did. A percentile calculator that does not state which population it is drawing from has not really given you a percentile at all, just a number shaped like one.
The same problem shows up at a smaller scale between two clinician-measured studies. Two studies can both be careful, both use trained observers, and still report somewhat different means, because they recruited from different clinics, different countries, or different age ranges - each a slightly different reference population, even before self-report enters the picture at all. A percentile is always relative to whichever one of these was used to build it, and swapping the reference without saying so changes the answer without changing anything about the man being measured - the figure moved on paper, not in reality. Age is one of the quieter variables in that recruitment gap - how age moves the published figures is worth checking before assuming a reference population matches you across every dimension.
What to check before trusting a percentile
Before treating a percentile as meaningful, confirm three things about its reference: whether it was clinician-measured or self-reported, what state it covers, and roughly how large and how it was recruited. The clinical protocol itself only gets you a trustworthy reading of your own figure - it says nothing about which reference population that figure should be checked against, and the two questions need answering separately rather than assuming a careful reading settles both.
A percentile from a service that will not name its reference population is not a number worth acting on, whatever confidence it is presented with, and the same caution applies further afield: a subjective score from Rate Cock, a rating tool, a human judge, or an AI image estimate is not a percentile of this kind at all, and none should be read as if it were one.