Method
Centimetres and scores are different objects
A length is a distance with an uncertainty; a rating is a judgement with a rubric. Confusing them is how a ruler ends up arguing with a score.
Two very different kinds of number get discussed as though they were the same thing. One is a length. The other is a judgement. Telling them apart clears up most of the confusion on this subject.
What makes something a measurement
A physical measurement has four properties, and all four are checkable by someone who was not in the room. It has a unit - centimetres, in the literature this site draws from. It has a defined procedure - bone-pressed, along the dorsal line, at a stated state - so that two people following the same instructions land on comparable numbers. It has an instrument with a known accuracy, a rigid ruler for length, a tape for girth, each with its own error bar. And it has an uncertainty: repeat the reading and it moves within a bounded range, which you can state explicitly rather than pretend away.
Take those four away and a figure stops being a measurement, whatever unit is printed after it.
What makes something a rating
A rating is built differently, on purpose. It has a rubric - a set of criteria, weighted somehow, that a judge or a model applies to what they are looking at. It has a judge, human or automated, whose training, mood, taste or calibration shapes the output in a way a ruler's calibration does not shape a length reading. And it is explicitly subjective: two competent judges looking at the same photograph can reasonably disagree, in a way two competent people bone-pressing the same shaft to the same bone should not.
None of that makes a rating worthless. It makes it a different category of information, answering a different question: not "how long is this" but "how does this compare, in someone's judgement, against others." Rate Cock is built around that second question - a score from a photograph, aggregated across raters - and it does not attempt to report a length in centimetres, because a photograph does not carry the scale information a length reading needs. A model doing the same work sits behind AI Penis; the aggregated version of that scoring, across many photos and raters, is what Penis Rater publishes; and a rating from an actual person rather than a model is what Rate Penis offers. All three of those are legitimate answers to "how does this look," which is a real question people have. None of them is an answer to "how long is this," which is a different question with a different method.
Where the two get tangled
The tangle usually runs one direction: someone takes a rating and treats it as though it implied a length, or takes a length and expects it to predict a rating. Neither inference holds up. A high rating does not mean a longer-than-average measurement, because a rating responds to proportion, presentation, angle and lighting as much as to raw size, and none of those are things a tape measures. A longer-than-average measurement does not guarantee a high rating either, for the same reasons run in reverse. The percentile a length occupies against published data and the score a photograph receives are computed from entirely different inputs, by entirely different processes, and there is no formula that converts reliably between them.
The other direction the tangle runs is treating a rating's number - out of ten, or a percentile-looking score - as though it carried the same kind of uncertainty a measurement does. It does not, because the source of the variation is different. A measurement's uncertainty comes from technique and instrument, and can be bounded by repeat readings. A rating's variation comes from disagreement between judges or between runs of a model, which is a property of the rubric and the raters, not of anything you did with a ruler.
Why the distinction is worth keeping straight
Practically, it means two things. If you want a number to compare against the published literature, take a measurement, using the standard method, and do not expect a rating from anywhere to substitute for it. If you want to know how something reads to another person, a rating answers that, and a ruler cannot, however carefully you use it. Asking a length to settle a question about judgement, or asking a judgement to settle a question about length, is the recurring mistake underneath a lot of arguments on this subject - and it dissolves the moment the two are kept as the separate objects they actually are.