Data

The studies have error bars too

Clinician measurement is better than self-measurement, not perfect; test-retest figures where reported set a floor under the precision of any pooled number.

3 min readData

The point of a clinician-measured study is that a trained observer with a rigid ruler removes most of what makes a self-measurement unreliable. It does not remove all of it. A health professional taking the same reading twice does not get the identical number twice, and that residual disagreement is worth naming rather than assuming away.

Two different questions

"Test-retest" and "inter-rater" sound similar and answer different questions. Test-retest asks: if the same observer measures the same man again, a few minutes or weeks apart, how close are the two readings? Inter-rater asks: if two different observers each measure the same man, how close are their readings to each other? This post is about the first. The second has its own literature and its own post, because the two do not have the same size or the same causes.

What drives test-retest disagreement

Even one trained observer with a consistent technique faces sources of variation that have nothing to do with skill. Erection state is never perfectly identical between two sessions. The pubic fat pad compresses by a slightly different amount depending on exactly how the ruler is seated that day. A tape's tension on a girth reading is a judgement made by hand, and a hand does not repeat itself to the tenth of a millimetre.

What the direction tells you

Studies that report a test-retest figure generally describe it as small relative to the between-person spread in the sample, which is the reassuring half of the finding: a clinician's single reading is a reasonable proxy for that person's "true" length or girth. It is not zero, though, and where a study reports it, it is typically on the order of a few millimetres rather than a fraction of one. That is a direction, not an invented number for any specific study - individual papers vary in how carefully they measured and reported this, and this post is not attaching a magnitude to a study that has not been checked.

What it does to the last decimal

A pooled mean like the erect length figure of 13.12 cm reported in the reference table of Veale figures by state is precise about the average across 15,521 men, because a large sample pins down a mean tightly regardless of how noisy any individual reading was. The test-retest figure is a different kind of precision: it bounds how much confidence to place in any one man's single measured reading, yours included. The practical consequence is that reading a published mean to one decimal place is appropriate, and reading your own single measurement to that same decimal place is not - the same caution belongs to a self-taken figure using the ordinary method, where the achievable precision is worse than a clinician's, not better.

Why this floor cannot be engineered away

It is tempting to assume better training or a stricter protocol could push test-retest disagreement toward zero, but the sources listed above are not procedural gaps - they are properties of a living body rather than of a sloppy observer. Erection state fluctuates by nature, the fat pad's compressibility is not a constant a protocol can fix, and hand tension on a tape will never be as repeatable as a machine's. A tighter protocol narrows the gap; it does not close it, which is why even the most careful clinical studies still report a nonzero test-retest figure rather than none at all.

Where this sits against everything else

Test-retest error inside a clinical study is the floor. Self-measurement error on top of a home ruler or tape sits above that floor, and where that additional error actually comes from is worth reading alongside this post rather than instead of it. None of this touches a different kind of number entirely - a rating drawn from a photograph is not attempting to be a measurement with an error bar in centimetres at all, it is a position in a distribution of judged images, assessed the way a scoring service assesses a photo or the way a human judge assesses one, and neither of those is trying to answer the question this post is about. An AI system estimating from an image has its own accuracy limits, which are a different and generally larger source of error than anything described here.

Read next

Full archive