Method
Where the error comes from
Two readings of the same person twenty minutes apart can differ by more than a centimetre. Almost none of that is the ruler.
Measure carefully three times in a row and you will get three different numbers. The spread is larger than people expect, and knowing where it comes from is what separates a measurement from a number you happened to write down.
The sources, largest first
Pressure at the pubic bone. By a distance the biggest one. The fat pad above the bone compresses under pressure, and how hard you press changes the reading directly. Between "resting the ruler against skin" and "pressed firmly to bone" there can be well over a centimetre, and the difference is entirely in the technique. This is the whole reason bone-pressed is the clinical standard.
Physiological state. Erection is not binary and the difference between "erect" and "fully erect" is real. Temperature, time since last activity, arousal level and time of day all move it. This source alone is why measuring once and treating it as definitive is optimistic.
Angle of the instrument. A ruler not parallel to the dorsal surface measures the hypotenuse of a triangle. Small angles cost little; a visible tilt costs several millimetres.
Reading and rounding. Parallax at an angle, plus the near-universal habit of rounding in the flattering direction. Small individually, and they compound with everything above.
The instrument itself. Almost nothing. A cheap ruler and an expensive one have the same millimetre markings. The instrument is the part people optimise and the part that matters least.
What precision is reachable
Careful technique, same session, three readings: roughly ±0.3 cm on length and ±0.2 cm on girth.
Across sessions on different days, the spread widens, because physiological state now varies too. Half a centimetre of session-to-session variation is unremarkable.
The rule that follows: a difference smaller than your error bar is not a difference. Two readings of 13.1 and 13.3 are the same reading. Treating them as a change is reading noise as signal, which is the most common error in this entire subject. The same rule applies to the variance in an AI score across retakes, which is larger than this and less under your control.
The protocol that removes most of it
- Same conditions each time - similar time of day, similar room temperature.
- Full erection, not partial.
- Bone-pressed, ruler along the top, parallel, read at eye level.
- Three readings. Take the median, not the best.
- Write down the method alongside the number.
Step four is the one people skip and it is doing the most work. The median of three is robust to a single bad reading in a way that a mean is not, and it removes the systematic upward bias you get from remembering the largest.
Step one is the same discipline that makes a rating submission comparable: fix the conditions, or the comparison is between setups.
Why this matters more than it sounds
Because the distribution is narrow - roughly the middle two thirds of erect lengths fall inside about a three-centimetre band. Against a spread that tight, a centimetre of measurement error is worth a large number of percentile points. Sloppy technique does not just give you a wrong number; it gives you a wrong number that lands somewhere quite different on the curve.
It is also a large part of why self-reported figures sit above measured ones. Error that is symmetric in the instrument becomes asymmetric in the reporting, because people remember their best reading.
Worth ending on the limit of the whole exercise: this gets you a repeatable figure in centimetres and nothing else. It says nothing about how anything is perceived, which is a separate question answered by a different kind of assessment entirely - the sort Rate Cock produces - and that judgement turns out to correlate with a tape measure far less closely than people assume. That goes double for a human reviewer's verdict, which is a response to a person and a presentation rather than to a figure.