Data
Two calculators, two answers
Online percentile calculators differ because they draw on different studies, different states and sometimes self-report; the disagreement is in the inputs, not the arithmetic.
Put the same erect length into two different percentile calculators and you can get two noticeably different percentiles back. The arithmetic behind a percentile calculation is simple and not where the disagreement comes from. The inputs are where it comes from, and most calculators do not show you theirs.
The reference study is the first variable
A percentile is only meaningful relative to a named population. A calculator built on Veale et al. (2015), BJU International, and its 15,521 clinically measured men will give a different answer than one built on an older, smaller study, or on a self-report survey with no clinical verification at all. Neither calculator's arithmetic is wrong once you accept its stated reference - they are simply answering the question against different populations, which is a different question wearing the same label.
If a calculator does not name its source study, there is no way to know which population your percentile is being measured against, which makes the output unverifiable regardless of how confident the interface looks.
State is the second variable, and it is often unlabelled
Erect, stretched flaccid and flaccid length have different means and different spreads. A calculator that does not ask which state you measured in is either assuming one silently or mixing figures from different states without saying so, and either way the percentile it returns may not correspond to the state you actually entered.
This matters more than it sounds, because the states are not simply offset by a constant. Stretched flaccid length tracks erect length reasonably well as a proxy, but flaccid length correlates with erect only weakly, so a calculator that quietly treats all three the same will be furthest wrong for flaccid inputs specifically.
Self-report versus clinical measurement is the third
Some calculators are built, deliberately or not, on self-reported reference data rather than clinically measured data. Self-reported figures run consistently higher than clinically measured ones, which means a calculator using a self-report reference will place an honest, carefully taken measurement at a lower percentile than a calculator built on Veale would. The measurement did not change. The population it is being compared against did.
How to tell which a calculator is using
Look for a named study and year in the calculator's fine print, the same way you would check a chart in a magazine article. Look for whether it asks which state you measured - erect specifically, not just "length" - before returning a number. And be suspicious of a calculator that returns a single precise-looking percentile with no stated source at all; that is the same provenance problem a redrawn press chart has, wearing an interactive interface instead of a printed page.
What a disagreement between two calculators actually tells you
It does not mean your measurement is unreliable. It means the two tools are answering against different reference populations, possibly in different states, and the disagreement is a property of the calculators rather than of you. The fix is the same one this site applies everywhere: trust the calculator that names its source and its state, and treat one that does not as unverified.
There is a fourth, quieter variable worth naming too: how the calculator handles uncertainty in your own input. Careful measurement still carries an error of a few millimetres, and a calculator that returns a single percentile to the nearest whole number, with no acknowledgement that the input itself has a margin, is presenting more precision than the underlying number supports - regardless of which reference study it used. A calculator that shows a range rather than a point estimate is doing something meaningfully more honest, even before you check which study sits behind it, and it is a small enough feature to build that its absence says something about how carefully the rest of the tool was built too.
A percentile calculator, however good, is still answering a different question than a subjective rating service like Rate Cock, a scoring hub's public results, or a human reviewer answers - a percentile locates a measurement in a distribution of measurements, and none of those is doing that. An AI tool estimating a figure from a photo is a fifth kind of number again, and feeding its output into any percentile calculator compounds two separate sources of error into one falsely precise result.